Regression

· by

Contents
A while ago, this link pointed to the content which is now in the Forecasting article.

Regression is one of the core tasks in machine learning. In this task, you get some input and your target variable is a single floating point number. For example, predicting the price of a house, estimating the age of the universe or calculating the probability that an image shows a dog. The age of the universe example shows that regression is not only used in machine learning and the dog image example shows that regression and classification can be very similar. A logistic regression can be converted to a classifier by choosing a threshold value (e.g. 0.5).

Big differences between regression and classification are the scoring functions and the targets. The targets in classification are just a few finite ones, while you have infinite possible targets for regression. Below, you can see a list of scoring functions.

Scoring functions

In the following, \(y\) is the ordered list of targets, \(y^P\) is the list of predictions in the same order and \(\bar{y}\) is the mean of \(y\).

Name Image X is better Definition and Usage
MAE $[0, \infty)$ lower $f(y, y^P) = \frac{1}{|y|} \sum_{y_i, y_i^P \in (y, y^P)} |y_i - y_i^P|$
MSE $[0, \infty)$ lower $f(y, y^P) = \frac{1}{|y|} \sum_{y_i, y_i^P \in (y, y^P)} (y_i - y_i^P)^2$
$R^2$ $(-\infty, 1]$ higher $f(y, y^P) = 1 - \frac{\sum (y_i - y_i^P)^2}{\sum (y_i - \bar{y})^2}$
Explained Variance $(-\infty, 1]$ higher $f(y, y^P) = 1 - \frac{Var(y - y^P)}{Var(y)}$

See also:

Regression Models

Trivial Models

There are some straightforward "models" for regression. They do learn, but they ignore the input completely:

  • Arithmetic mean: \(\frac{1}{n}\sum_{i=1}^n {y_i}\)
  • Median: Sort all \(y_i\) and take the value in the middle
  • minimum and maximum
  • q-Quantile: Sort the \(y_i\) and take the first value after going through \(q \in [0, 1]\) of the input. For \(q = 0.5\), this is the median.
  • Other "means" like the geometric mean

Linear regression

Linear regression tries to fit a line to the input by minimizing the squared distance between the input points and the line. This usually gives pretty good results.

The model looks like this:

$$\hat{y}(x) = \sum_{i=1}^n b_i \cdot x_i \text{ with }b_i \in \mathbb{R}$$

If one defines \(X \in \mathbb{R}^{m \times n}\) (one row per sample), one can also write it in a vectorized form:

$$\hat{y}(X) = X \cdot \beta \text{ with }\beta \in \mathbb{R}^n$$

It can be "learned" (calculated) with

$$\beta = {(X^T X)}^{-1} X^T y$$

Example: Fitting a line with least squares

To see where such a formula comes from, look at the simplest case. Suppose you have \(n\) 2D data points \((x_1, y_1), (x_2, y_2), \dots, (x_n, y_n)\) and you want the line \(y = a \cdot x + b\) that fits those data points best.

Now we have to think about what "fits best" means. The least squares method finds \(a\) and \(b\) that minimize the sum of squared errors \(f: \mathbb{R}^2 \rightarrow \mathbb{R}\) with

$$f(a, b) := \sum_{i=1}^n \left ((a \cdot x_i + b) - y_i \right )^2$$

Let's say \(a\) was constant. Then \(f\) is a quadratic function in \(b\), and we find its minimum where the derivative is zero:

$$\begin{align} f(a,b) &= \sum_{i=1}^n \left ((a \cdot x_i + b) - y_i \right )^2\\ &= \sum_{i=1}^n ((a x_i +b)^2 - 2 (a x_i+b) y_i + y_i^2)\\ &= \sum_{i=1}^n ((a x_i)^2 + 2 a x_i b + b^2 -2 (a x_i+b) y_i + y_i^2)\\ \frac{\partial f(a,b)}{\partial b} &= \sum_{i=1}^n (2a x_i +2b - 2y_i)\\ 0 &\stackrel{!}{=} 2 \sum_{i=1}^n (a x_i +b - y_i)\\ \Leftrightarrow 0 &\stackrel{!}{=} a \sum_{i=1}^n x_i + n \cdot b - \sum_{i=1}^n y_i\\ \Leftrightarrow b &\stackrel{!}{=} \frac{\sum_{i=1}^n y_i -a \sum_{i=1}^n x_i}{n}\\ \Leftrightarrow b &\stackrel{!}{=} \bar y -a \bar x \end{align}$$

where \((\bar x, \bar y)\) is the center of all points. So the best line always goes through the center of the points. Let's do the same with \(a\) and insert \(b = \bar y - a \bar x\) and \(\sum_{i=1}^n x_i = n \bar x\):

$$\begin{align} \frac{\partial f(a,b)}{\partial a} &= \sum_{i=1}^n (2a x_i^2 +2 x_i b - 2 x_i y_i)\\ 0 &\stackrel{!}{=} a \sum_{i=1}^n x_i^2 + b \sum_{i=1}^n x_i - \sum_{i=1}^n x_i y_i\\ \Leftrightarrow 0 &\stackrel{!}{=} a \sum_{i=1}^n x_i^2 + (\bar y - a \bar x) \cdot n \bar x - \sum_{i=1}^n x_i y_i\\ \Leftrightarrow a &\stackrel{!}{=} \frac{\sum_{i=1}^n x_i y_i - n \bar x \bar y}{\sum_{i=1}^n x_i^2 - n \bar x^2} = \frac{\sum_{i=1}^n (x_i - \bar x)(y_i - \bar y)}{\sum_{i=1}^n (x_i - \bar x)^2} \end{align}$$

This only works if not all \(x_i\) are equal; otherwise the denominator is zero and every line through the center is equally good. As \(f\) is a sum of squares, this point is a minimum and not a maximum. It is the same result as the matrix formula above with

$$X = \begin{pmatrix}x_1 & 1\\ \vdots & \vdots\\ x_n & 1\end{pmatrix} \text{ and } \beta = \begin{pmatrix}a\\ b\end{pmatrix}$$

Logistic regression

Logistic regression tries to fit the logistic function \(y(x) = \frac{1}{1+e^{-x}}\) to the input. This function is nice as it is within \([0, 1]\) and thus can be used to represent a probability.

Trees

You can also use trees for regression. One idea of how to do that is by "bucketing" observations and applying one of the trivial models to each bucket. Such models can only predict values between what they observed before.

See also: