Log in Sign up
Back to Discover
🔢

Prior probability

math Maturity 11-13

You can guess things before they happen. You might guess if it will be hot tomorrow. You use what you know to help you. Then you see what really happens. This helps you learn more. What do you think will happen next?

42 words

You can make a guess before you see new facts. This guess is called a prior. It is what you think is true right now.

You might guess the weather for tomorrow. You can use what happened today to help. This makes your guess better.

Sometimes you have no idea at all. You might guess that every choice is just as likely. This is a very simple guess.

When you see new facts, you change your guess. You mix your old guess with the new news. This helps you learn.

Your new guess is based on what you just saw. It is how we learn about the world.

108 words

Imagine you are guessing the weather for tomorrow. You might look at today's temperature to help. This starting guess is called a prior probability. It is what you believe before you see new facts.

Some priors give a lot of detail. These are called informative priors. For example, you might guess tomorrow will be warm based on past years. Other priors give very little detail. These are called uninformative priors. You might use these when you have no idea at all. One way to do this is to say every choice is just as likely. If a ball is under one of three cups, you might guess each cup has a one-third chance.

When you see new data, you update your guess. This new guess is called the posterior probability. You mix your old prior with the new news. If your first guess was very strong, the new data might not change it much. If you have no info, the new data will change your guess a lot. This is a key part of Bayesian statistics. It is how we use math to learn from the world.

187 words

Imagine you are trying to guess the temperature at noon tomorrow. You might start by looking at the temperature today. This initial guess is called a prior probability. It is what you assume before you see any new evidence. In math, a prior describes how likely different outcomes are before we collect data. This helps us build a starting point for learning. We can use priors to model many things, like how people might vote in an election.

There are different ways to build these starting guesses. An informative prior uses specific details to guide us. For example, you could use the temperature from today to guess tomorrow. A weakly informative prior gives only a little bit of information. It helps keep our guesses in a reasonable range without being too strict. If we have no information at all, we use an uninformative prior. This is sometimes called an objective prior. One simple rule for this is the principle of indifference. This means we assign equal chances to every possible choice.

People have studied these ideas for a long time. Some mathematicians believe priors should be based on personal opinions. These are called subjective Bayesians. Others believe we can find priors that are logically required. These are called objective Bayesians. Edwin T. Jaynes was a famous thinker in this group. He used ideas like symmetry to explain why certain guesses make sense. For instance, if a ball is under one of three cups, you might guess each cup has a one-third chance. This is because swapping the labels of the cups should not change your guess.

Other experts have also made important discoveries about priors. J.B.S. Haldane proposed a special type of prior in 1932. This is known as the Haldane prior. It focuses heavily on cases where something happens every time or never at all. Another mathematician named Harold Jeffreys created a systematic way to design uninformative priors. He developed the Jeffreys prior for certain types of math problems. In modern times, computers use special methods called Markov chain Monte Carlo to handle these complex calculations. These tools make it much easier to work with difficult math.

Learning from a prior is a way to connect what we know to what we see. When we get new data, we use Bayes' rule to update our guess. This new, updated guess is called the posterior probability. If our first guess was very strong, the new data might not change it much. If our first guess was weak, the new data will change it a lot. This process is how we turn old assumptions into new knowledge. It shows how math helps us constantly learn from the world around us.

454 words

A prior probability distribution represents an assumed set of probabilities for an uncertain quantity before any new evidence is considered. In Bayesian statistics, this serves as a starting point for reasoning about the world. The uncertain quantity might be a parameter within a mathematical model or a latent variable that cannot be observed directly. For example, a prior might represent the expected proportions of voters for a specific politician before election results are tallied. Once new data is collected, mathematicians use Bayes' rule to update this prior. This process produces a posterior probability distribution, which is the updated conditional distribution of the quantity given the new evidence.

There are several ways to construct a prior distribution depending on the information available. An informative prior expresses specific, definite information about a variable. For instance, if you are predicting the temperature at noon tomorrow, you might use a normal distribution centered on today's noon temperature. A strong prior is a specific type of informative prior where the initial assumption is so dominant that new data barely changes it. Conversely, a weakly informative prior provides only partial information. It steers an analysis toward reasonable solutions without being too restrictive. This is often used for regularization, which keeps statistical inferences within a sensible range.

When no information is available, researchers may use an uninformative prior. These are also called objective or diffuse priors. One classic method for creating these is the principle of indifference, which assigns equal probabilities to all possible outcomes. In many parameter estimation problems, using an uninformative prior leads to results similar to conventional statistical analysis. This happens because the likelihood function from the data often provides more information than the vague prior. Some mathematicians, known as objective Bayesians, believe that certain priors are logically required by the nature of uncertainty. They argue that these priors can be found through mathematical principles rather than personal opinion.

Edwin T. Jaynes was a major figure in the objective Bayesian school. He used the concept of symmetry to justify certain priors. For example, if a ball is hidden under one of three cups labeled A, B, and C, the most logical prior is a uniform distribution where each cup has a 1/3 probability. Jaynes argued that if you swap the labels, your prediction should not change. This invariance principle suggests that the uniform prior is the only choice that preserves the symmetry of the problem. He also championed the principle of maximum entropy, or MAXENT. This principle suggests choosing the distribution that contains the least amount of information consistent with known constraints.

Other mathematicians have developed specialized priors for specific mathematical needs. In 1932, J.B.S. Haldane proposed the Haldane prior. This is an improper prior distribution that places extreme weight on the possibility that an event will either always happen or never happen. If a researcher observes a chemical dissolving in one experiment but not in another, applying Bayes' theorem to the Haldane prior results in a uniform distribution. Harold Jeffreys also contributed by creating a systematic way to design uninformative priors, such as the Jeffreys prior for Bernoulli random variables. These mathematical tools allow for more precise handling of different types of data and uncertainty.

Advanced modeling often involves hierarchical priors. In these systems, the prior distributions of model parameters depend on their own parameters, known as hyperparameters. For example, if you use a beta distribution to model a parameter $p$ of a Bernoulli distribution, the $\alpha$ and $\beta$ values of the beta distribution are the hyperparameters. This allows for many conditional levels of distributions within a single model. Historically, choosing priors was difficult because mathematicians had to use conjugate families to ensure the posterior remained easy to calculate. However, the widespread availability of Markov chain Monte Carlo methods has made this constraint much less of a concern in modern science.

Modern applications also use priors for mechanical purposes like feature selection. Researchers like José-Miguel Bernardo introduced reference priors to maximize the information gained from data. This involves maximizing the expected Kullback–Leibler divergence between the prior and the posterior. By doing this, the researcher ensures the prior is as "least informative" as possible about the parameter. Whether through subjective judgment or objective mathematical principles, the prior remains a fundamental component of how we bridge the gap between what we assume and what we observe.

728 words
Up Next
🔢
Posterior probability
Math
More to explore

What is Nepedia?

A free, ad-free encyclopedia for children. Every article is written at five reading levels, so the same page works for a five-year-old and a fifteen-year-old — use the level switcher above to see this one change. No account needed to read.