Machine Learning Summary Section Two
This post was written in 2016 and may be outdated.
Linear Regression with Multiple Variables
- $h_{\theta}(x)=\theta_{0} + \theta_{1}x_{1} + \theta_{2}x_{2} + … + \theta_{n}x_{n}$
- $x_{0} = 1$ , $h_{\theta}(x)=\theta_{0}x_{0} + \theta_{1}x_{1} + \theta_{2}x_{2} + … + \theta_{n}x_{n}$
- $h_{\theta}(x)=\theta^{T}X$
Feature Scaling
- make sure features are on a similar scale
- get every feature into approximately a $-1\leq x_{i} \leq +1$ range
- mean normalization $ x_{i} = \frac{x_{i} - \mu_{i}}{s_{i}} $ . ($ \mu_{i} $ is avg and $ s_{i} $ can be $ max - min $ or standard deviation )
Learning Rate
- making sure gradient descent is working correctly. $ J(\theta) $ should decrease after every iteration.
- $ J(\theta)-iteration $ curve goes up or oscillates, lower the learning rate. The learning rate should be moderate — too small makes gradient descent painfully slow
- To try $ \alpha $, try … 0.001, 0.003, 0.01, 0.03, 0.1, 0.3, 1, 3 …
Features and Polynomial Regression
Linear regression can’t fit all data; we need other models
Normal Equation
- For some linear regression problems, we can use the normal equation. Complexity: $O(n^{3})$
- Training-set feature matrix $X$ (with $x_{0} = 1$); training results as vector y
- Solve for the vector with the normal equation:\(\theta = (X^{T}X)^{-1}X^{T}y\)
- $X^{T}X$ forms a square matrix, which is what makes it invertible.$X\theta = y$
Octave or Matlab
- Basic operations
- Moving data
- Computing on data
- Plotting data
- Control statements and functions
- Vectorization
Logistic Regression
- Hypothesis of the logistic regression model: $h_{\theta}(x) = g(\theta_{T}X)$
- Sigmoid function: $ g(z) = \frac{1}{1 + e^{-z}} $
- Cost function of logistic regression: \(J(\theta) = \frac{1}{m} \sum_{i = 1}^{m}Cost(h_\theta(x^{(i)}),y^{(i)})\)
- Simplified: \(Cost(h_\theta(x),y) = -y \times log(h_\theta(x)) - (1-y) \times log(1 - h_\theta(x))\)
- fminunc
- one-vs-all: take $max(h_\theta(x))$ as the classification result
Regularization
- Overfitting — high variance
- Solutions to overfitting
- Reduce the number of features (manually select which features to keep, or use a model selection algorithm)
- Regularization
- Applying regularization to linear and logistic regression: $\lambda$