Linear Algebra · Undergraduate

Least Squares and the Line of Best Fit

Quick answer

An inconsistent system Ax = b has no exact solution, but it has a best approximate one. A least-squares solution x̂ makes Ax̂ the point of the column space closest to b, which is the orthogonal projection of b onto Col A. The error b − Ax̂ is then orthogonal to every column of A, which is the statement Aᵀ(b − Ax̂) = 0, or the normal equations AᵀAx̂ = Aᵀb. Fitting a line y = β₀ + β₁x to data points is the most common case, and it minimizes the sum of the squared vertical errors.

What you'll learn

  • Explain a least-squares solution geometrically
  • Set up and solve the normal equations
  • Fit a line of best fit to data with a design matrix
  • Check a fit through its residuals

When there is no solution

Four measurements, one line: the points (1,1)(1, 1), (2,3)(2, 3), (3,4)(3, 4) and (4,4)(4, 4) do not lie on any line y=β0+β1xy = \beta_0 + \beta_1x. Demanding that they did gives four equations in two unknowns,

[11121314][β0β1]=[1344]\begin{bmatrix} 1 & 1\\ 1 & 2\\ 1 & 3\\ 1 & 4 \end{bmatrix} \begin{bmatrix} \beta_0 \\ \beta_1 \end{bmatrix} = \begin{bmatrix} 1 \\ 3 \\ 4 \\ 4 \end{bmatrix}

and the system is inconsistent. The useful question is not whether it has a solution but which β\boldsymbol{\beta} comes closest.

A least-squares solution of Ax=bA\mathbf{x} = \mathbf{b} is a vector x^\hat{\mathbf{x}} with ∥b−Ax^∥≤∥b−Ax∥\|\mathbf{b} - A\hat{\mathbf{x}}\| \le \|\mathbf{b} - A\mathbf{x}\| for every x\mathbf{x}.

The name comes from the quantity minimized: ∥b−Ax∥2\|\mathbf{b} - A\mathbf{x}\|^2 is the sum of the squares of the individual errors.

Why the normal equations find it

As x\mathbf{x} varies, AxA\mathbf{x} runs over the column space of AA. The point of Col A\text{Col}\,A closest to b\mathbf{b} is its orthogonal projection b^\hat{\mathbf{b}}, and the error b−b^\mathbf{b} - \hat{\mathbf{b}} is orthogonal to Col A\text{Col}\,A. So x^\hat{\mathbf{x}} is exactly a solution of Ax=b^A\mathbf{x} = \hat{\mathbf{b}}, and the error must be orthogonal to every column aj\mathbf{a}_j:

aj⋅(b−Ax^)=0for every j\mathbf{a}_j \cdot (\mathbf{b} - A\hat{\mathbf{x}}) = 0 \quad\text{for every } j

The transpose ATA^{\mathsf{T}} has the columns of AA as its rows, so ATvA^{\mathsf{T}}\mathbf{v} lists the dot products of the columns with v\mathbf{v}. All the orthogonality conditions together read AT(b−Ax^)=0A^{\mathsf{T}}(\mathbf{b} - A\hat{\mathbf{x}}) = \mathbf{0}, which rearranges to the normal equations:

ATA x^=ATbA^{\mathsf{T}}A\,\hat{\mathbf{x}} = A^{\mathsf{T}}\mathbf{b}

The best approximation leaves an error perpendicular to the column space, and the normal equations are that perpendicularity written column by column. When the columns of AA are independent, ATAA^{\mathsf{T}}A is invertible and the solution is unique.

The line of best fit

For data (x1,y1),…,(xn,yn)(x_1, y_1), \ldots, (x_n, y_n) and a line y=β0+β1xy = \beta_0 + \beta_1x, build the design matrix XX from a column of 11s and the column of xx-values, and solve

XTX β=XTyX^{\mathsf{T}}X\,\boldsymbol{\beta} = X^{\mathsf{T}}\mathbf{y}

The errors being squared are the vertical distances from the points to the line, called residuals.

Worked examples

Common mistakes

Practice problems

  1. Find the transpose of [123456]\begin{bmatrix} 1 & 2 & 3\\ 4 & 5 & 6 \end{bmatrix}.

    Answer

    [142536]\begin{bmatrix} 1 & 4\\ 2 & 5\\ 3 & 6 \end{bmatrix}

    Full solution

    The rows become the columns.

  2. For the points (1,1)(1, 1), (2,2)(2, 2) and (3,2)(3, 2), write the design matrix and compute XTXX^{\mathsf{T}}X and XTyX^{\mathsf{T}}\mathbf{y}.

    Answer

    XTX=[36614]X^{\mathsf{T}}X = \begin{bmatrix} 3 & 6\\ 6 & 14 \end{bmatrix} and XTy=(5,11)X^{\mathsf{T}}\mathbf{y} = (5, 11)

    Full solution

    XX has rows (1,1)(1, 1), (1,2)(1, 2) and (1,3)(1, 3). The entries of XTXX^{\mathsf{T}}X are 33 points, 1+2+3=61 + 2 + 3 = 6, and 1+4+9=141 + 4 + 9 = 14. And XTy=(1+2+2, 1+4+6)X^{\mathsf{T}}\mathbf{y} = (1 + 2 + 2,\ 1 + 4 + 6).

  3. Solve the normal equations from exercise 2 for the best line.

    Answer

    y=23+12xy = \tfrac{2}{3} + \tfrac{1}{2}x

    Full solution

    3β0+6β1=53\beta_0 + 6\beta_1 = 5 and 6β0+14β1=116\beta_0 + 14\beta_1 = 11. Doubling the first and subtracting gives 2β1=12\beta_1 = 1, so β1=12\beta_1 = \tfrac{1}{2} and β0=5−33=23\beta_0 = \tfrac{5 - 3}{3} = \tfrac{2}{3}.

  4. Find the least-squares line for (−1,0)(-1, 0), (0,1)(0, 1) and (1,3)(1, 3).

    Answer

    y=43+1.5xy = \tfrac{4}{3} + 1.5x

    Full solution

    XTX=[3002]X^{\mathsf{T}}X = \begin{bmatrix} 3 & 0\\ 0 & 2 \end{bmatrix} because the xx-values add to 00, and XTy=(4,3)X^{\mathsf{T}}\mathbf{y} = (4, 3). So β0=43\beta_0 = \tfrac{4}{3} and β1=32\beta_1 = \tfrac{3}{2}.

  5. Find the residuals of the fit in exercise 4, and check that they add to zero.

    Answer

    16\tfrac{1}{6}, −13-\tfrac{1}{3} and 16\tfrac{1}{6}

    Full solution

    The line gives −16-\tfrac{1}{6}, 43\tfrac{4}{3} and 176\tfrac{17}{6} at the three xx-values. Subtracting from 00, 11, 33 gives the residuals, and 16−26+16=0\tfrac{1}{6} - \tfrac{2}{6} + \tfrac{1}{6} = 0.

  6. Find the least-squares solution of [111−1]x=[31]\begin{bmatrix} 1 & 1\\ 1 & -1 \end{bmatrix}\mathbf{x} = \begin{bmatrix} 3 \\ 1 \end{bmatrix}.

    Answer

    (2,1)(2, 1), the exact solution

    Full solution

    The system is consistent, so the smallest possible error, zero, is reached at its solution. The normal equations [2002]x^=[42]\begin{bmatrix} 2 & 0\\ 0 & 2 \end{bmatrix}\hat{\mathbf{x}} = \begin{bmatrix} 4 \\ 2 \end{bmatrix} agree.

  7. When does Ax=bA\mathbf{x} = \mathbf{b} have exactly one least-squares solution?

    Answer

    When the columns of AA are linearly independent

    Full solution

    Then ATAA^{\mathsf{T}}A is invertible and x^=(ATA)−1ATb\hat{\mathbf{x}} = (A^{\mathsf{T}}A)^{-1}A^{\mathsf{T}}\mathbf{b}. With dependent columns, many x\mathbf{x} give the same projection Ax^A\hat{\mathbf{x}}.

  8. What is Ax^A\hat{\mathbf{x}}, geometrically?

    Answer

    The orthogonal projection of b\mathbf{b} onto Col A\text{Col}\,A

    Full solution

    Ax^A\hat{\mathbf{x}} is the point of the column space closest to b\mathbf{b}, and the closest point of a subspace is the orthogonal projection.

  9. If b\mathbf{b} is orthogonal to every column of AA, and the columns are independent, what is x^\hat{\mathbf{x}}?

    Answer

    x^=0\hat{\mathbf{x}} = \mathbf{0}

    Full solution

    Then ATb=0A^{\mathsf{T}}\mathbf{b} = \mathbf{0}, and the normal equations ATAx^=0A^{\mathsf{T}}A\hat{\mathbf{x}} = \mathbf{0} with invertible ATAA^{\mathsf{T}}A force x^=0\hat{\mathbf{x}} = \mathbf{0}. The projection of b\mathbf{b} onto the column space is the origin.

  10. A student fits a line to four points by picking two of them and drawing the line through those. What went wrong?

    Hint

    What happens to the other two points?

    Answer

    The line ignores half the data. Least squares uses all four points and minimizes the total squared error.

    Full solution

    The line through (2,3)(2, 3) and (4,4)(4, 4), for instance, is y=2+0.5xy = 2 + 0.5x, with residuals −1.5-1.5, 00, 0.50.5 and 00 and squared error 2.52.5, worse than the 11 achieved by y=0.5+xy = 0.5 + x. Only the normal equations balance every point at once.

Frequently asked questions

What is a least-squares solution?

A vector x̂ that makes ‖b − Ax̂‖ as small as possible. When Ax = b is consistent it is an ordinary solution; otherwise it is the best approximation.

What are the normal equations?

AᵀAx̂ = Aᵀb. They say the error b − Ax̂ is orthogonal to every column of A.

Why is it called least squares?

It minimizes ‖b − Ax‖², which is the sum of the squares of the individual errors.

When is the least-squares solution unique?

When the columns of A are linearly independent. Then AᵀA is invertible and x̂ = (AᵀA)⁻¹Aᵀb.

How do you fit a line to data with matrices?

Put a column of 1s and the column of x-values into a design matrix X, the y-values into a vector y, and solve XᵀXβ = Xᵀy for the intercept and slope.