1 Basic definitions

1.1 Signed measures and probability measures

Let \((\Omega,\mathcal F)\) be a measurable space. A signed measure \(\mu\) on \((\Omega,\mathcal F)\) is a function that assigns a real number to each measurable set and is countably additive, allowing negative values. A probability measure \(P\) is a special case: it is nonnegative and satisfies \(P(\Omega)=1\).

Signed measures arise naturally when subtracting two probability measures (e.g., \(P-Q\)). Many definitions related to total variation are therefore formulated for signed measures first, then specialized to probability measures.

1.2 Total variation of a signed measure

For a signed measure \(\mu\), its total variation \(\mu\) is the unique measure characterized by

\[

\mu(A)=\sup\left\{\sum_{i=1}^n\mu(A_i):\; A_1,\dots,A_n \in \mathcal F \text{ partition } A\right\}.

\]

Intuitively, \(\mu\) measures the maximal accumulated magnitude of \(\mu\) across partitions of sets.
When \(\mu\) has a Jordan decomposition \(\mu=\mu^+-\mu^-\) (with \(\mu^+,\mu^-\) nonnegative and mutually singular), then \(\mu=\mu^+ + \mu^-\).

1.3 Total variation norm (distance form)

The total variation norm of a signed measure \(\mu\) is \[

\|\mu\|_{\mathrm{TV}} :=\mu(\Omega).

\] For two probability measures \(P\) and \(Q\), their difference \(P-Q\) is a signed measure, and the total variation distance is commonly taken as \[

d_{\mathrm{TV}}(P,Q):=\|P-Q\|_{\mathrm{TV}}.

\] Some authors insert a factor \(1/2\) to match a probability-of-error convention; this is discussed in the variants section.

1.4 Equivalent characterizations

1.4.1 Supremum over measurable events

A fundamental characterization expresses total variation distance using events: \[

\|P-Q\|_{\mathrm{TV}}=\sup_{A\in\mathcal F}P(A)-Q(A).

\] Thus, two distributions are close in total variation precisely when every measurable event has nearly the same probability under both.

1.4.2 Supremum over bounded test functions

Another equivalent form uses bounded measurable functions. For measurable \(f\) with \(\|f\|_\infty\le 1\),

\[

\|P-Q\|_{\mathrm{TV}}=\sup_{\|f\|_\infty\le 1}\left\int f\,d(P-Q)\right.

\] This viewpoint treats total variation as the strongest possible bias one can detect using any bounded “test” function.

1.4.3 Discrete/finite-space reduction

If \(\Omega\) is finite or countable, and \(P,Q\) admit mass functions \(p(x),q(x)\), then \[

\|P-Q\|_{\mathrm{TV}}=\frac12\sum_{x\in\Omega}p(x)-q(x)

\] under the standard convention where the distance is half the \(\ell^1\) distance. Correspondingly, without the half-factor convention, the formula scales accordingly. In either case, the computation reduces to summing pointwise discrepancies.

1.4.4 Density-based formula (absolutely continuous case)

If \(P\) and \(Q\) are absolutely continuous with respect to a common dominating measure \(\lambda\), with densities \(p=\frac{dP}{d\lambda}\) and \(q=\frac{dQ}{d\lambda}\), then \[

\|P-Q\|_{\mathrm{TV}}=\frac12\int_\Omegap-q\,d\lambda

\] (with the same convention remarks as in the discrete case). This expresses the distance as the integrated separation between the two density profiles.

2 Properties and intuition

2.1 Metric properties and bounds

The total variation distance between probability measures is a metric: it is nonnegative, symmetric, satisfies the triangle inequality, and vanishes exactly when the measures agree.

Moreover, it is bounded above. Under the common \(1/2\)-adjusted convention, \(0\le d_{\mathrm{TV}}(P,Q)\le 1\). Under the unadjusted norm convention described earlier, the upper bound scales accordingly, but the qualitative interpretation remains: distances are finite and capped.

2.2 Range for probability measures

For probability measures \(P,Q\), the signed measure \(P-Q\) has total variation controlled by the fact that both are normalized to unit mass. Consequently, the total variation distance cannot exceed a maximal value corresponding to complete separation (events that one distribution assigns probability one to and the other assigns probability zero to).

2.3 Interpretation via distinguishability

Total variation measures how well one can distinguish two distributions using observations of a single random draw. Large total variation means there exist events that differ substantially in probability, so outcomes provide strong evidence favoring one distribution over the other. Small total variation means all events behave similarly under both models.

This “uniform over events” nature is why total variation is often regarded as a strong notion of closeness.

2.4 Relationship to coupling

2.4.1 Coupling inequality

A central link is the coupling inequality: there exists a coupling \((X,Y)\) with marginals \(P\) and \(Q\) such that \[

\mathbb P(X\ne Y)\le \|P-Q\|_{\mathrm{TV}}.

\] Conversely, the best possible coupling (minimizing \(\mathbb P(X\ne Y)\)) achieves this value under the standard conventions. Couplings therefore translate total variation distance into an explicit probability of mismatch.

2.5 Triangle inequality and stability

The triangle inequality implies stability under intermediate approximations: if \(P\) is close to \(Q\) and \(Q\) is close to \(R\), then \(P\) must be close to \(R\). This property underlies many approximation arguments, where distributions are built in stages and errors are accumulated additively.

3 Connections to convergence

3.1 Total variation convergence

A sequence of probability measures \((P_n)\) converges in total variation to \(P\) if \[

\|P_n-P\|_{\mathrm{TV}}\to 0.

\] This is a very strong mode of convergence. It implies convergence of probabilities for every measurable event: \[

\sup_{A\in\mathcal F}P_n(A)-P(A)\to 0.

\] Hence, no event becomes a persistent discrepancy in the limit.

3.2 Implications for other modes of convergence

Total variation convergence implies weaker forms such as convergence in distribution, because convergence of probabilities for sets in a suitable generating class transfers to distributional convergence. It also implies convergence in many integral metrics for bounded functions due to the test-function characterization.

However, the reverse is generally false: weak convergence can occur while total variation distance stays bounded away from zero, particularly when distributions remain mutually singular or change on increasingly fine structures.

Total variation convergence is compatible with strong control of integrals. For families of functions, especially if they are uniformly bounded, the test-function characterization gives immediate convergence of expectations. When unbounded functions are involved, additional assumptions such as uniform integrability become relevant to justify passing to limits.

In practice, one uses total variation for bounded observables and supplements it with integrability conditions for unbounded ones.

3.4 Concentration of probability mass

Because total variation captures discrepancies over all events, it forces probability mass to concentrate similarly under both measures. If \(P_n\) puts substantial mass on a set \(A_n\) while \(P\) assigns nearly none to that set (or vice versa), the distance cannot be small. Therefore, convergence in total variation prevents “mass from drifting” into regions that the limit measure views as negligible.

4 Statistical testing viewpoint

4.1 Hypothesis testing and minimal error probability

In a binary hypothesis test distinguishing \(P\) from \(Q\) based on a single observation, total variation governs the minimal achievable error when priors are balanced and the decision rule is optimal.

At a high level, the supremum characterization over events identifies the decision boundary that maximizes the gap \(P(A)-Q(A)\), yielding the best discrimination. Thus, total variation acts as an operational quantity: it equals the best possible advantage over random guessing, up to the conventional factor \(1/2\).

4.2 Likelihood ratio perspective (high-level)

When densities exist, testing can be expressed in terms of the likelihood ratio \(p/q\). Optimal tests often compare \(p\) and \(q\) pointwise and reject one hypothesis when the likelihood ratio crosses a threshold (Neyman–Pearson-type logic). While the likelihood ratio framework depends on model structure, total variation remains the global summary of how distinguishable the two distributions are.

4.3 Decision rules achieving the bound

The event-supremum form suggests constructing a set \(A^\star\) such that \(P(A^\star)-Q(A^\star)\) is maximal. A corresponding rule—deciding for \(P\) if the observation lies in \(A^\star\), and for \(Q\) otherwise—achieves the extremal performance tied to the total variation distance.

5 Computation and estimation

5.1 Finite sample spaces: algorithms and formulas

On finite spaces, computation reduces to evaluating the \(\ell^1\) discrepancy between mass functions. If one has empirical estimates \(\hat p(x)\) and \(\hat q(x)\), then total variation can be computed by summing \(\hat p(x)-\hat q(x)\) over all states (with the relevant \(1/2\) convention). In settings with sparse support, this can be done efficiently by focusing only on points where either estimate is nonzero.

5.2 Lower bounds and upper bounds

When exact computation is infeasible, one can bound total variation using partial information: bounds derived from a restricted class of events or test functions provide provable guarantees. In density settings, upper bounds can be obtained from approximations of \(p-q\) in \(L^1\), while lower bounds can be obtained by finding a candidate set \(A\) with large probability gap.

5.3 Estimating total variation in practice (conceptual)

5.3.1 Empirical measures

Given samples from \(P\) and \(Q\), a naive plug-in approach estimates \(d_{\mathrm{TV}}(P,Q)\) by replacing \(P,Q\) with their empirical measures. The quality of such estimates depends on sample size, dimension, and how rapidly the empirical mass stabilizes. The resulting procedure is straightforward but can require many samples when the distributions differ only on small-probability regions.

5.3.2 Kernel or smoothing approaches (overview)

In continuous settings, estimation often uses smoothing or density estimation: kernels can approximate densities, then the \(L^1\) distance between estimated densities is computed (again with appropriate scaling conventions). The effectiveness depends on bandwidth choices, regularity assumptions, and how estimation error propagates into the absolute value in the \(L^1\) integrand.

6 Applications in probability

6.1 Markov chains and mixing times

6.1.1 Mixing in total variation

For Markov chains, total variation distance provides a direct measure of how quickly the chain approaches its stationary distribution. If \(P^t(x,\cdot)\) denotes the distribution after \(t\) steps starting from \(x\), then the chain is said to mix when \(\|P^t(x,\cdot)-\pi\|_{\mathrm{TV}}\) becomes small, where \(\pi\) is the stationary measure. This metric yields mixing-time bounds that are interpretable in terms of worst-case distinguishability from equilibrium.

6.2 Marked point processes (general idea)

In models where events are indexed by time and an additional mark (type, size, category), distributions can often be compared through total variation of the induced laws on the space of marked configurations. While direct total variation computation can be complex, the concept is used to assess whether two generative mechanisms produce statistically similar patterns of marked events.

6.3 Randomized algorithms and distributional approximation

Many randomized algorithms rely on replacing an ideal distribution with a computable surrogate. Total variation distance is a standard yardstick for quantifying how much this substitution changes observable behavior. Because it controls probabilities of all events, it gives robust guarantees: any downstream decision depending on the output distribution is affected only in proportion to the total variation discrepancy.

6.4 Robustness and sensitivity of distributions

The strength of total variation lies in its uniform event control. In sensitivity analyses, it formalizes the idea that small perturbations in a model should not drastically alter observable outcomes. When total variation is small, one can bound changes in likelihood of any event and in expectations of bounded observables.

7.1 Half total variation distance convention

Some texts define the half total variation distance as \[

\mathrm{TV}(P,Q)=\frac12\|P-Q\|_{\mathrm{TV}},

\] so that it equals the maximal difference over events in a convention aligned with testing probabilities. This factor harmonizes formulas with expressions like \(\mathbb P(X\ne Y)\) under optimal coupling and with minimal error probability in simple hypothesis testing settings. The underlying concept is unchanged; only scaling conventions differ.

7.2 Hellinger distance comparison (overview)

The Hellinger distance is another measure of divergence between probability distributions, often expressed via square roots of densities. It satisfies metric properties up to conventions and is connected to affinity between distributions. Compared with total variation, Hellinger distance is frequently more tractable in estimation and offers different continuity behavior.

7.3 Kullback–Leibler divergence comparison (overview)

The Kullback–Leibler (KL) divergence quantifies expected log-likelihood ratio under one distribution. Unlike total variation, KL is not a metric and can be infinite when one distribution is not absolutely continuous with respect to the other. KL captures average information discrepancy rather than worst-case event differences, and it relates to total variation through inequalities under suitable conditions.

7.4 Wasserstein distance contrast (overview)

The Wasserstein distance compares distributions by considering the cost of transporting mass over a metric space. Total variation ignores geometry and focuses purely on how probability mass is allocated to events. Wasserstein distance can be sensitive to the “distance” between regions where mass shifts, making it useful when spatial structure matters.

7.5 Norms on measures beyond total variation

Beyond total variation, one can consider other norms on signed measures, such as norms induced by different test-function classes or by integral operators. These alternatives trade off sensitivity to support mismatches, smoothness requirements, and computational tractability, while still providing quantitative comparison between probability laws.