1 Definition and basic properties
A probability measure is a function that assigns a number between 0 and 1 to events in a set of possible outcomes. It provides the formal language for discussing chance, uncertainty, and random variation. In modern mathematics, probability measures are used to define random variables, distributions, expectations, and limit theorems.
1.1 Measure-theoretic formulation
In measure theory, a probability measure is a measure whose total mass is 1. It is defined on a sigma-algebra of subsets of a set, rather than on arbitrary subsets. This framework allows one to handle both finite and infinite sample spaces in a precise way. The assignment of probabilities is required to be consistent under countable unions of disjoint events.
1.2 Axioms of a probability measure
A probability measure satisfies three central axioms that determine its behavior. These axioms are sufficient to derive most standard rules of probability. They ensure that probability remains compatible with set operations and limit processes.
1.2.1 Nonnegativity
Every event is assigned a probability that is greater than or equal to zero. Negative likelihoods are excluded, since they would not admit a sensible interpretation as chance. This axiom also implies that the measure of the empty set is zero.
1.2.2 Normalization
The entire sample space has probability one. This expresses the idea that, when one of the permitted outcomes is realized, something in the sample space must occur. Normalization distinguishes probability measures from more general measures, which may have any finite or infinite total mass.
1.2.3 Countable additivity
If a collection of events is pairwise disjoint, then the probability of their union equals the sum of their probabilities. This property extends to countably many disjoint events. Countable additivity is essential for treating limiting processes, infinite outcome spaces, and events built from sequences of simpler events.
1.3 Probability space
A probability space is the basic setting in which probability theory is formulated. It consists of a sample space, a sigma-algebra of events, and a probability measure. Together, these objects determine which events are measurable and how likely they are.
1.3.1 Sample space
The sample space is the set of all possible outcomes of a random experiment. Its elements are often called outcomes or elementary events. Depending on the problem, the sample space may be finite, countably infinite, or uncountable.
1.3.2 Sigma-algebra
A sigma-algebra is a collection of subsets of the sample space that is closed under complements and countable unions. It specifies the events to which probabilities can be assigned. The restriction to measurable sets avoids pathologies that can arise when one tries to measure every subset of an uncountable space.
1.3.3 Probability measure
The probability measure assigns a numerical probability to each event in the sigma-algebra. It encodes the statistical structure of the model and determines the likelihood of composite events. In applications, different probability measures can represent different assumptions about randomness.
2 Examples
Probability measures appear in many familiar forms. Some are discrete, some are continuous, and others combine both behaviors. Examples help show how the abstract axioms correspond to concrete models.
2.1 Finite probability spaces
In a finite sample space, a probability measure is often given by assigning weights to each outcome. The probability of an event is the sum of the weights of its elements. If all outcomes are equally likely, each receives the same probability.
2.2 Discrete distributions
A discrete distribution assigns positive probability to countably many outcomes, with all other outcomes having probability zero. Common examples include the Bernoulli, binomial, and geometric distributions. The probability measure is determined by a probability mass function.
2.3 Continuous distributions
Continuous distributions spread probability across intervals rather than individual points. A single point typically has probability zero, while intervals may have positive probability. The normal distribution and the uniform distribution on an interval are standard examples.
2.4 Mixed distributions
Some probability measures have both discrete and continuous parts. For instance, a distribution may place positive mass at specific points while also having a continuous component elsewhere. Such models are useful when data include both fixed values and variable measurements.
3 Properties and consequences
The axioms of probability imply many useful rules for manipulating events. These rules are frequently used in calculations and proofs. They also connect set-theoretic relations with numerical inequalities.
3.1 Complements and monotonicity
The probability of a complement equals one minus the probability of the original event. If one event is contained in another, then its probability cannot be larger. These facts follow directly from additivity and normalization.
3.2 Finite additivity
For finitely many disjoint events, the probability of their union is the sum of their probabilities. This is a direct consequence of countable additivity. Finite additivity is often used in elementary probability calculations.
3.3 Continuity of probability
Probability measures behave continuously with respect to increasing or decreasing sequences of events. This is one of the most important features of measure-theoretic probability. It allows one to pass to limits in many constructions.
3.3.1 Continuity from below
If events increase to a limiting event, then their probabilities increase to the probability of the limit. This property is called continuity from below. It is especially useful when approximating complicated events by simpler ones.
3.3.2 Continuity from above
If events decrease to a limiting event and the initial event has finite probability, then their probabilities decrease to the probability of the limit. For probability measures, the finiteness condition is automatically satisfied. This makes decreasing approximations especially convenient.
3.4 Inclusion-exclusion principle
The inclusion-exclusion principle gives formulas for the probability of unions of events in terms of intersections. For two events, it subtracts the overlap once to avoid double counting. For more events, alternating sums correct for repeated overlap among several sets.
4 Construction of probability measures
Probability measures are often built indirectly from simpler ingredients. The construction may begin with values on a limited collection of sets and then extend to a larger sigma-algebra. Such methods are central in rigorous probability theory.
4.1 Defining measures on simple spaces
On simple spaces, such as finite or countable sets, a probability measure can be specified by listing point probabilities. On intervals or geometric regions, it may be defined by densities or by symmetry considerations. These initial definitions are typically chosen so that the total mass is one.
4.2 Extension theorems
Extension theorems provide conditions under which a premeasure defined on a smaller family of sets extends to a full probability measure. They are crucial when constructing measures on complicated spaces from consistent local data. Many standard probability models rely on these results.
4.2.1 Carathéodory extension theorem
Carathéodory’s extension theorem states that a premeasure on an algebra can be extended uniquely to a measure on the generated sigma-algebra, under suitable conditions. This theorem is a fundamental tool in measure theory. It is often used to build probability measures from finite-dimensional specifications.
4.2.2 Kolmogorov extension theorem
The Kolmogorov extension theorem gives conditions for constructing a probability measure on an infinite product space from a consistent family of finite-dimensional distributions. It is a key result in stochastic process theory. The theorem ensures that compatible finite descriptions determine a global process.
4.3 Probability measures on product spaces
Product spaces arise when considering multiple random variables or repeated experiments. A product probability measure describes the joint behavior of several components. Independence is often expressed by requiring the measure of a product set to factor into the measures of the individual sets.
5 Probability measures on common spaces
Certain measure spaces appear so often that they serve as standard models. They provide intuition and are widely used in analysis, statistics, and probability theory. Each has characteristic properties and applications.
5.1 Counting measure and normalized counting probability
Counting measure assigns to each set the number of its elements, when this number is finite or countably infinite. On a finite set, normalizing the counting measure produces the uniform probability measure. Each outcome then has equal weight.
5.2 Lebesgue measure and uniform probability
Lebesgue measure is the standard notion of length, area, or volume on Euclidean spaces. When restricted and normalized on a bounded region, it yields a uniform probability distribution. This is the usual formalization of “equally likely” points in a continuum.
5.3 Gaussian measures
Gaussian measures are probability measures associated with normal distributions. On the real line, they are determined by a mean and variance. In higher-dimensional settings, Gaussian measures are central in analysis, statistics, and stochastic processes.
5.4 Borel probability measures
A Borel probability measure is defined on the Borel sigma-algebra of a topological space. These measures are particularly important on Euclidean spaces and other topological settings. They allow one to connect probability with continuity, convergence, and topology.
6 Random variables and distributions
Random variables translate events into numerical quantities. Their distributions are probability measures induced on target spaces such as the real numbers. This perspective is one of the main advantages of measure-theoretic probability.
6.1 Pushforward measures
The distribution of a random variable is the pushforward of the underlying probability measure under that variable. This means that probabilities of sets in the target space are computed by taking preimages in the original space. Pushforward measures provide a general way to describe distributions without relying on formulas.
6.2 Distribution functions
For real-valued random variables, the distribution function gives the probability that the variable is less than or equal to a given value. It is nondecreasing, right-continuous, and ranges from 0 to 1. Distribution functions encode the essential features of one-dimensional laws.
6.3 Joint distributions
Joint distributions describe the combined behavior of several random variables. They are probability measures on product spaces and capture dependence among variables. Marginal distributions are obtained by ignoring some coordinates.
6.4 Conditional probability measures
Conditional probability measures describe probabilities after information has been taken into account. They refine the original measure relative to a given event or sigma-algebra. Such measures are used to model updating, dependence, and sequential observation.
7 Expectation and integration
Expectation is defined using integration with respect to a probability measure. This unifies sums, averages, and weighted values in a single framework. Many statistical quantities are expressed naturally in this language.
7.1 Integrals with respect to a probability measure
An integral with respect to a probability measure averages a function over the sample space according to its probabilities. For simple random variables, the integral is a weighted sum; for general functions, it is obtained by approximation and limit processes. This approach is more flexible than classical summation formulas.
7.2 Expected value
The expected value of a random variable is its average value under the probability measure. It summarizes the central tendency of a distribution. Expected value can be finite, infinite, or undefined, depending on the integrability of the random variable.
7.3 Moments and variance
Moments are expectations of powers of a random variable. They describe features such as location, spread, skewness, and tail behavior. Variance measures dispersion around the mean and is defined as the expectation of the squared deviation from that mean.
7.4 Almost sure concepts
An event occurs almost surely if it has probability one. Almost sure statements are stronger than statements that are merely highly likely, yet they may still allow exceptional cases of probability zero. This notion is widely used in limit theorems and stochastic analysis.
8 Convergence and limit theorems
Probability measures support several distinct notions of convergence. These notions are essential for describing how random quantities behave in the long run. Limit theorems formalize the emergence of stable patterns from repeated random trials.
8.1 Convergence of events
Sequences of events may converge by increasing toward a limit, decreasing toward a limit, or through more general set-theoretic behavior. Probability continuity allows the probabilities of such sequences to converge as well. This is often used in proving asymptotic results.
8.2 Modes of convergence
Random variables can converge in several different senses, each capturing a different type of approximation. These modes are related but not equivalent. Choosing the appropriate one depends on the application.
8.2.1 Almost sure convergence
Almost sure convergence means that, except on a set of probability zero, the sequence converges pointwise. It is one of the strongest common forms of convergence in probability theory. It is especially useful when one wants pathwise behavior.
8.2.2 Convergence in probability
Convergence in probability means that the chance of a large deviation from the limit becomes small. This is weaker than almost sure convergence but still strong enough for many statistical arguments. It expresses approximation in a probabilistic rather than pointwise sense.
8.2.3 Convergence in distribution
Convergence in distribution concerns the convergence of distribution functions or, more generally, of induced measures. It is weaker than convergence in probability and often appears in asymptotic approximation. This mode is central in limit theorems for sums of random variables.
8.3 Laws of large numbers
The laws of large numbers describe the stabilization of sample averages as the number of observations grows. They explain why repeated random trials tend to produce predictable aggregate behavior. These results are foundational in statistics and applied probability.
8.4 Central limit theorem
The central limit theorem states that suitably normalized sums of many independent random variables tend toward a Gaussian law. It is one of the most celebrated results in probability. The theorem explains the frequent appearance of normal distributions in natural and statistical phenomena.
9 Advanced topics
Advanced probability theory studies finer structural properties of measures and their relationships. These ideas are essential in modern analysis, stochastic processes, and statistical inference. They extend the basic framework to more sophisticated settings.
9.1 Absolute continuity and singularity
One probability measure is absolutely continuous with respect to another if every null set of the second is also null for the first. Two measures are singular if they concentrate on essentially disjoint sets. These notions classify how measures overlap or differ.
9.2 Radon-Nikodym derivatives
The Radon-Nikodym derivative describes the density of one measure relative to another when absolute continuity holds. In probability, it often appears as a likelihood ratio or change-of-measure factor. It provides a systematic way to transform expectations and probabilities.
9.3 Conditional expectation
Conditional expectation is the best approximation of a random variable given partial information, in an average-squared sense. It is itself a random variable measurable with respect to the available information. This concept generalizes ordinary averaging and underlies much of martingale theory.
9.4 Probability measures on infinite-dimensional spaces
Probability measures on infinite-dimensional spaces arise in the study of stochastic processes, random fields, and functional data. These spaces may consist of paths, functions, or sequences. Constructing and analyzing measures there often requires careful topological and measure-theoretic tools.