Think about a coin flip. 
Imagine you flip a coin. 
Imagine you ask a question with only two answers. You might ask if it will rain today. The answer is either yes or no. This is a single trial. A mathematician named Jacob Bernoulli studied these kinds of events. 
We call this the Bernoulli distribution. It is a way to model a yes or no choice. In math, we often use the number 1 for success. We use the number 0 for failure. We use a letter, p, to show the chance of success. For example, a coin might have a chance to land on heads. If the coin is fair, the chance is half. But some coins are unfair. They might land on heads more often.
This idea is a part of the binomial distribution. A binomial distribution looks at many trials. A Bernoulli distribution only looks at one. It also helps us measure uncertainty. We call this measure entropy. Uncertainty is highest when both outcomes are equally likely. If one outcome is certain, the entropy is zero. This helps us understand randomness in our world.
Imagine you are asking a question that only has two possible answers. You might ask if a coin will land on heads or tails. You could also ask if a light bulb will turn on or stay off. In math, we call this a single trial. This simple idea helps us model many things in the real world. We use a special way to track these results called a Bernoulli distribution. 
To make this work, we assign numbers to our two choices. We use the number 1 to represent a success or a "yes" answer. We use the number 0 to represent a failure or a "no" answer. We use the letter p to show the probability of success. The chance of failure is called q. For example, if a coin is unfair, p might be higher for heads. This lets us describe how likely each result is.
A mathematician named Jacob Bernoulli studied these ideas. He was from Switzerland and helped shape how we see math. The Bernoulli distribution is a special case of other math ideas too. If you do the same trial many times, it becomes a binomial distribution. In a binomial distribution, the number of trials is called n. For a Bernoulli distribution, n is always exactly 1. This makes it the simplest building block for more complex math. 
Mathematicians use this to measure things like uncertainty and information. They call the measure of uncertainty "entropy." Entropy is highest when both outcomes are equally likely. If you know for sure what will happen, the entropy is zero. They also use "Fisher information" to measure how much data we have. This helps us learn about unknown things by watching trials.
You can see this math in many places every day. It is used in a "Bernoulli process" which is a long string of these trials. It also links to the geometric distribution. That describes how many tries you need until you finally get a success. Even the way computers use bits relates to this. A bit is just a 1 or a 0. 
The Bernoulli distribution is a fundamental concept in probability theory and statistics. It is a discrete probability distribution for a random variable. This variable can only take two possible values: 1 or 0. In many models, the value 1 represents a success or a "yes" outcome. The value 0 represents a failure or a "no" outcome. This distribution acts as a model for any single experiment that asks a yes-no question. Such questions result in Boolean-valued outcomes. These are single bits where the value is success, yes, true, or one with a probability of p. The probability of failure, or zero, is represented by q. 
To understand the mechanism, we look at the probability mass function. This function describes the probability for each possible outcome k. The probability of success is p, and the probability of failure is q. Because there are only two outcomes, p and q must add up to 1. You can express this relationship in several ways. One way is to say the probability is p for k=1 and q for k=0. Another way is to use the formula p^k * q^(1-k). This mathematical structure allows us to calculate the likelihood of a specific result in a single trial.
The Bernoulli distribution is a specific type of several broader distributions. It is a special case of the binomial distribution where the number of trials, n, is exactly 1. It is also a special case of a two-point distribution. In a two-point distribution, the outcomes do not have to be 0 and 1. Furthermore, the Bernoulli distributions for different values of p form what is known as an exponential family. If you perform many independent and identical Bernoulli trials, their sum follows a binomial distribution. This connects simple single events to larger sets of data. 
History connects this idea to the Swiss mathematician Jacob Bernoulli. His work helped define how we understand random variables and chance. Mathematicians have since used his namesake distribution to develop complex statistical tools. For example, the maximum likelihood estimator for p, based on a random sample, is simply the sample mean. This means we can estimate the true probability of success by looking at the average of our results. This connection between theory and observation is a cornerstone of modern statistics.
We can use specific numbers to describe the properties of this distribution. The expected value, or the mean, of a Bernoulli random variable is p. The variance, which measures how much the results spread out, is calculated as p times q. Because p and q are probabilities, the variance will always stay between 0 and 1. The skewness, which describes the asymmetry of the distribution, is calculated using the formula (q - p) divided by the square root of pq. These precise values allow scientists to predict the behavior of random systems with great accuracy.
Mathematicians also use the Bernoulli distribution to study uncertainty and information. Entropy is a measure used to describe the randomness in a distribution. For a Bernoulli variable, entropy is maximized when p is 0.5. This means there is the highest level of uncertainty when both outcomes are equally likely. If p is 0 or 1, the entropy is zero because the outcome is certain. Another important concept is Fisher information. This measures how much information an observable random variable carries about an unknown parameter p. Fisher information is also maximized when p is 0.5, reflecting maximum uncertainty. 
Finally, the Bernoulli distribution relates to many other important mathematical ideas. The geometric distribution models how many independent Bernoulli trials are needed to reach one success. The categorical distribution is a generalization used for variables with any constant number of discrete values. There is also the Beta distribution, which serves as the conjugate prior of the Bernoulli distribution. In some cases, if p is 0.5, the distribution is known as a Rademacher distribution. These connections show how one simple yes-no model can support a vast network of mathematical thought.
🖼️ Images & Media (2)
More to explore
✨ What else?
Related topics you might enjoy
🔬 Go deeper
More advanced topics to explore
🪜 Step back
Simpler topics to build understanding
What is Nepedia?
A free, ad-free encyclopedia for children. Every article is written at five reading levels, so the same page works for a five-year-old and a fifteen-year-old — use the level switcher above to see this one change. No account needed to read.