We use dots to show what we see.
Sometimes we collect many facts. These facts are like dots on a page.
One way is to use a math rule. This rule finds a path that stays very close to all the dots. It tries to make the gaps small.
Many smart people studied this. A man named Gauss used it to find a space rock. 
It is a very useful tool for science. It helps us see patterns in the world.
When we collect facts, they often look like dots on a page.
Many smart people studied this idea. In 1805, Legendre wrote a clear guide on how to use it. A man named Gauss also worked on it. He used this math to find a space rock called Ceres. 
Imagine you are looking at a scatter of dots on a graph. These dots represent real things we have measured, like the height of plants or the speed of a car.
To make this work, we look at how much each dot misses our line. We take each gap, or residual, and square it. Squaring means multiplying the number by itself. We do this so that negative gaps and positive gaps do not cancel each other out. Then, we add all those squared numbers together. The goal is to find the specific path where this total sum is the smallest it can be.
Many famous thinkers helped develop this idea over hundreds of years. In 1700, Isaac Newton used a method of averages to study the equinoxes. Later, in 1750, Tobias Mayer used similar ideas to study the Moon. In 1757, Roger Joseph Boscovich used it to study the shape of the Earth. Pierre-Simon Laplace also worked on this problem in 1788 and 1789. He tried to find a way to minimize errors using a specific mathematical form. These early steps paved the way for the formal math we use today.
In 1805, Adrien-Marie Legendre published the first clear guide for this method. He showed how to use algebra to fit lines to data. Shortly after, Carl Friedrich Gauss published his own work in 1809. Gauss claimed he had used the method since 1795. He went even further by connecting least squares to the normal distribution. 
This math is incredibly useful for predicting the future. A great example happened in 1801 with a new space rock called Ceres.
Least squares is a mathematical method used to find the best-fit model for a set of data. When we collect observations, the data points rarely form a perfect line or curve. Instead, they scatter around a central trend due to measurement errors or random fluctuations. The least squares method aims to adjust a model's parameters to minimize the total error. This error is measured by the sum of the squared residuals. A residual is the vertical distance between an observed data point and the value predicted by the model.
To understand the mechanism, we must look at how residuals are treated. For every data point, we calculate the difference between the actual observation and the model's prediction. We then square each of these differences. Squaring is essential because it ensures that positive and negative deviations do not cancel each other out. By summing these squared values, we create a single number representing the total error. The "best" model is the one that makes this total sum as small as possible. This process is often achieved by setting the gradient of the error function to zero.
Least squares problems generally fall into two distinct categories. The first is linear least squares, also known as ordinary least squares. In these problems, the model is a linear combination of its unknown parameters. These problems are highly efficient because they have a closed-form solution, meaning we can calculate the answer directly using matrix algebra. The second category is nonlinear least squares. In these cases, the model functions are not linear in all unknowns. Because there is usually no direct formula for nonlinear models, we must use iterative refinement. This means we start with an initial guess and repeatedly improve it through successive approximations.
The history of this method is a culmination of eighteenth-century mathematical advances. Isaac Newton used a method of averages to study equinoxes in 1700. He even wrote down early versions of the normal equations used in ordinary least squares. In 1750, Tobias Mayer applied similar ideas to the Moon's librations. Later, Roger Joseph Boscovich used the method of least absolute deviation to study the Earth's shape in 1757. Pierre-Simon Laplace also worked on these problems in 1788 and 1789. Laplace attempted to define a method of estimation that minimized error by using a symmetric two-sided exponential distribution. 
The formalization of the method arrived in the early nineteenth century. In 1805, Adrien-Marie Legendre published the first clear exposition of least squares. He described it as an algebraic procedure for fitting linear equations to data. Shortly after, in 1809, Carl Friedrich Gauss published his work on celestial orbits. Gauss claimed he had used the method since 1795, which led to a priority dispute with Legendre. However, Gauss provided a deeper connection between least squares and the principles of probability. He showed that the arithmetic mean is the best estimate when errors follow a normal distribution.
A famous demonstration of Gauss's work involved the asteroid Ceres. Giuseppe Piazzi discovered Ceres on 1 January 1801, but it was lost behind the Sun after 40 days. Astronomers needed to predict its location without solving extremely complex nonlinear equations. The 24-year-old Gauss used least-squares analysis to succeed where others could not. His predictions allowed Franz Xaver von Zach to relocate the asteroid. This success proved the immense power of the method in the field of astronomy. In 1822, Gauss further established that his approach was optimal for linear models with specific error conditions. This result is part of the Gauss–Markov theorem.
While powerful, the method has specific limitations regarding how it handles error. Standard least squares assumes that errors only exist in the dependent variable. It implicitly assumes that errors in the independent variable are zero or negligible. If errors exist in both variables, researchers may use a different approach called total least squares. This method balances the effects of different error sources to create a more pragmatic fit. Understanding these distinctions is vital when using regression for prediction versus fitting a "true relationship."
🖼️ Images & Media (6)
More to explore
✨ What else?
Related topics you might enjoy
🔬 Go deeper
More advanced topics to explore
🪜 Step back
Simpler topics to build understanding
What is Nepedia?
A free, ad-free encyclopedia for children. Every article is written at five reading levels, so the same page works for a five-year-old and a fifteen-year-old — use the level switcher above to see this one change. No account needed to read.