1 Definition and basic properties
A pushforward measure formalizes how a measure on a source measurable space is transported to a target space via a measurable map. It is the basic tool for turning “events” in the domain into corresponding events in the codomain by applying a function.
1.1 Measurable spaces and measurability requirements
Let \( (X,\Sigma_X) \) and \( (Y,\Sigma_Y) \) be measurable spaces. A map \(f:X\to Y\) is required to be measurable, meaning that for every \(B\in\Sigma_Y\), the preimage \(f^{-1}(B)\) lies in \(\Sigma_X\). Let \(\mu\) be a measure on \((X,\Sigma_X)\).
1.2 Definition via preimages
Given a measurable function \(f:X\to Y\) and a measure \(\mu\) on \(X\), the pushforward measure \(f_\*\mu\) on \(Y\) is defined by \[ (f_\*\mu)(B)=\mu\!\left(f^{-1}(B)\right), \quad B\in\Sigma_Y. \] This definition assigns to each measurable set in the target the \(\mu\)-mass of all points in \(X\) that map into it.
1.2.1 Pushforward of measures on measurable rectangles (when applicable)
When \(X\) and \(Y\) are product spaces and \(f\) has a form that interacts well with the product \(\sigma\)-algebra, the pushforward can often be computed on generating sets such as measurable rectangles. For example, for maps that act coordinatewise or collapse one factor, evaluating \((f_\*\mu)(A\times C)\) typically reduces to \(\mu\) of corresponding preimages built from rectangles.
1.3 Relationship with probability measures
If \(\mu\) is a probability measure on \(X\), then \(f_\*\mu\) is a probability measure on \(Y\). In probability theory, this construction is exactly the way one obtains the distribution of a transformed random variable: if \(X\) is the sample space, \(\mu\) is the law of an \(X\)-valued random element, and \(f\) is the transformation, then \(f_\*\mu\) is the law of \(f\) applied to that random element.
1.4 Identity, composition, and invariance checks
Two basic sanity checks are standard.
- Identity map: If \(\mathrm{id}_X:X\to X\) is the identity, then \((\mathrm{id}_X)_\*\mu=\mu\).
- Composition: If \(f:X\to Y\) and \(g:Y\to Z\) are measurable, then \( (g\circ f)_\*\mu = g_\*(f_\*\mu)\). This property ensures pushforwards behave consistently under successive transformations.
2 Existence, well-definedness, and measurability
Although the definition appears simple, several technical points ensure it produces a genuine measure on the target \(\sigma\)-algebra.
2.1 Conditions ensuring the target is measurable
The preimage rule is the key: for \(f_\*\mu\) to be well-defined on \(\Sigma_Y\), it is necessary that every set \(B\in\Sigma_Y\) has a measurable preimage \(f^{-1}(B)\in\Sigma_X\). No further structure is required; once measurability holds, \(B\mapsto \mu(f^{-1}(B))\) automatically defines a measure on \(Y\).
2.2 Handling null sets under pushforwards
Pushforwards respect measure-zero sets in a particular direction: if \(A\subseteq X\) satisfies \(\mu(A)=0\), then any target set \(B\) whose preimage lies in \(A\) will satisfy \((f_\*\mu)(B)=0\). However, the converse need not hold: a set in \(Y\) may have zero pushforward mass even if its preimage has complicated structure on \(X\).
2.2.1 Measures concentrated on sets of measure zero
If \(\mu\) is concentrated on a \(\mu\)-null set from a competing measure perspective (for instance, when comparing two different measures), the pushforward inherits the same concentration behavior relative to the target measure constructed from \(\mu\). This is often used when transferring absolute continuity or singularity questions through a map.
2.3 Functorial viewpoint (maps of measure spaces)
In categorical language, a measurable map between spaces induces a map between measures by pushforward. This perspective highlights that the pushforward construction depends only on the measurable structure and the measure itself.
2.3.1 Compatibility with composition \( (g\circ f)_\* \)
The relation \((g\circ f)_\*\mu = g_\*(f_\*\mu)\) is the hallmark of functorial behavior. It means that pushing forward in one step yields the same result as pushing forward in stages, which is essential for iterated transformations and layered models.
3 Fundamental measure-theoretic behavior
Pushforwards preserve many order and size properties of measures, making them compatible with standard measure-theoretic operations.
3.1 Linearity and monotonicity
If \(\mu\) and \(\nu\) are measures on \(X\) and \(a,b\ge 0\), then \[ f_\*(a\mu+b\nu)=a\,f_\*\mu+b\,f_\*\nu. \] Moreover, monotonicity holds: if \(\mu\le \nu\) (meaning \(\mu(B)\le \nu(B)\) for all measurable \(B\subseteq X\)), then \(f_\*\mu \le f_\*\nu\).
3.2 Total mass and finiteness
Total mass is preserved by pushforward: \[ (f_\*\mu)(Y)=\mu(f^{-1}(Y))=\mu(X). \] Consequently, if \(\mu\) is finite or \(\sigma\)-finite, these properties pass to the pushforward in the natural way.
3.3 Preservation of \(\sigma\)-finiteness
If \(X\) can be covered by countably many measurable sets \(X_n\) with \(\mu(X_n)<\infty\), then \(f(X_n)\) need not be measurable as a set in a naive sense, but one can work instead with preimages of measurable sets. A standard argument yields that \(f_\*\mu\) is \(\sigma\)-finite whenever \(\mu\) is.
3.4 Behavior under restrictions to measurable subsets
If \(A\subseteq X\) is measurable and \(\mu_A\) denotes the restriction \(\mu_A(B)=\mu(A\cap B)\), then the pushforward of \(\mu_A\) is closely related to the pushforward of \(\mu\), concentrated on the image of \(A\) as seen through preimages.
3.4.1 Pushforward of restricted measures
More precisely, \[ f_\*(\mu_A)(B)=\mu_A(f^{-1}(B))=\mu(A\cap f^{-1}(B)). \] This equals \((f_\*\mu)(B\cap f(A))\) in contexts where \(f(A)\) is measurable or when interpreting it via measurable hulls. In full generality, the reliable statement is in terms of \(A\cap f^{-1}(B)\), which is always measurable.
4 Interaction with measure classes and decompositions
Pushforward interacts strongly with the structure used to compare measures: absolute continuity, singularity, and densities.
4.1 Absolute continuity under pushforward
A classical theme is: if one measure is absolutely continuous with respect to another on the domain, then the same relationship often holds for their pushforwards on the target. Concretely, if \(\nu\ll \mu\) on \(X\), then \(f_\*\nu \ll f_\*\mu\) on \(Y\). The proof is direct: for any measurable \(B\subseteq Y\), if \((f_\*\mu)(B)=0\) then \(\mu(f^{-1}(B))=0\), so \(\nu(f^{-1}(B))=0\), hence \((f_\*\nu)(B)=0\).
4.2 Singularity considerations
Singularity can behave more subtly. Even if \(\nu\) and \(\mu\) are singular on \(X\), their pushforwards may fail to be singular on \(Y\) when the map collapses disjoint sets into overlapping regions. The pushforward can “merge” mass from different parts of \(X\), reducing the distinguishability of the measures.
4.3 Radon–Nikodym derivatives of pushforwards
When \(f_\*\nu\ll f_\*\mu\), a Radon–Nikodym derivative exists: there is a measurable function \(h\) on \(Y\) such that \[ \frac{d(f_\*\nu)}{d(f_\*\mu)}=h. \] Identifying \(h\) in terms of densities on \(X\) typically requires additional structure, such as an explicit formula for conditional expectations or a regularity assumption on the map and spaces.
4.3.1 Characterizing densities in special settings
In special cases—such as smooth maps between Euclidean spaces with measures having densities—\(h\) can often be expressed via Jacobians and conditional weighting along fibers. In more abstract settings, \(h\) is frequently characterized through identities of the form \[ \int_Y \varphi(y)\,d(f_\*\nu)(y)=\int_Y \varphi(y)\,h(y)\,d(f_\*\mu)(y), \] for test functions \(\varphi\), and then \(h\) is pinned down by the Radon–Nikodym theorem.
5 Integration and change-of-variables
Pushforwards provide a clean framework for change-of-variables: rather than integrating over the domain with complicated constraints, one integrates over the codomain against the pushed measure.
5.1 Integral transformation formula
For any measurable function \(\varphi:Y\to\mathbb{R}\) for which the integrals make sense, \[ \int_Y \varphi(y)\,d(f_\*\mu)(y)=\int_X \varphi(f(x))\,d\mu(x). \] This identity is the core computational bridge between integrals on \(X\) and integrals on \(Y\).
5.1.1 From simple functions to general measurable functions
The formula is often established first for simple functions \(\varphi\) (finite linear combinations of indicator functions), using the definition of \(f_\*\mu\). Extension to nonnegative measurable \(\varphi\) follows by monotone convergence, and to integrable \(\varphi\) by decomposing into positive and negative parts or applying dominated convergence under standard hypotheses.
5.2 Expectation of transformed random variables
When \(\mu\) is a probability measure, the identity becomes \[ \mathbb{E}[\varphi(f(X))]=\int_Y \varphi(y)\,d(f_\*\mu)(y), \] so expectations of functions of transformed variables are computed using the law \(f_\*\mu\).
5.3 Change of variables in one dimension (typical smooth case)
| In one-dimensional smooth settings, if \(f:\mathbb{R}\to\mathbb{R}\) is differentiable and monotone on relevant intervals, and if \(\mu\) has a density \(p\) with respect to Lebesgue measure, then the density of the pushforward can be expressed using the derivative of \(f\). The essential pattern is that the mapping rescales intervals according to \( | f'(x) | \). |
|---|
5.3.1 Jacobian factors and orientation conventions
| Orientation conventions matter when using determinants in more general settings. In one dimension, the absolute value \( | f'(x) | \) ensures the pushforward density remains nonnegative; whether one writes \( | f' | \) or splits by monotonicity intervals depends on how the inverse map is used. |
|---|
5.4 Multidimensional transformations and conditions
In \(\mathbb{R}^n\), similar formulas involve Jacobian determinants. However, not every measurable map supports a neat density expression; additional regularity (such as differentiability and appropriate injectivity or local invertibility) is typically required.
5.4.1 The role of measurable bijections and inverses
When \(f\) is a measurable bijection with a measurable inverse and suitable differentiability properties, the transformation of densities becomes straightforward. The inverse allows rewriting integrals over \(Y\) as integrals over \(X\), yielding explicit Jacobian factors that track how volume elements change under the map.
6 Pushforward and pullback relationships
Pushforward on measures and pullback on observables form a dual pair: measures move forward, while functions move backward.
6.1 Pullback of functions and induced operators
Given a measurable \(f:X\to Y\), the pullback of a function \(\psi:Y\to\mathbb{R}\) is \(\psi\circ f:X\to\mathbb{R}\). In integral terms, this pullback is exactly what appears on the right-hand side of the transformation formula: \[ \int_X (\psi\circ f)\,d\mu. \] Thus pushforward on measures can be viewed as the adjoint operation to pullback on functions.
6.2 Duality between pushforward on measures and pullback on observables
The defining identity \[ \int_Y \psi\,d(f_\*\mu)=\int_X (\psi\circ f)\,d\mu \] encapsulates the duality: testing a measure on \(Y\) against \(\psi\) is equivalent to testing \(\mu\) on \(X\) against \(\psi\circ f\).
6.3 Transport identities
Transport identities describe how expectations, integrals, and related quantities commute with the mapping through pushforward.
6.3.1 Commuting integrals under mapping
Whenever integrals are well-defined and measurability conditions hold, one can “move the map” from the argument of the function to the measure. This is often used to simplify calculations in probability and analysis by transferring complexity from the integrand to the measure—or vice versa.
7 Disintegration perspective
Disintegration refines pushforward by decomposing a measure into conditional measures along the fibers of a map.
7.1 Intuition: conditioning along fibers
Given \(f:X\to Y\), points in \(Y\) correspond to subsets (fibers) of \(X\) of the form \(f^{-1}(\{y\})\). Disintegration expresses the original measure \(\mu\) as an “average” of conditional measures living on these fibers, weighted by the pushforward measure \(f_\*\mu\).
7.2 Existence of disintegration under standard hypotheses
Exact existence depends on structural assumptions on the spaces and measures (commonly standard Borel spaces and \(\sigma\)-finite measures). Under such hypotheses, one can find a family of measures \(\{\mu_y\}_{y\in Y}\) on \(X\) concentrated on the fibers, together with the property that their average reconstructs \(\mu\).
7.3 Reconstructing the original measure from conditional measures
A typical reconstruction statement is that for suitable test functions \(\Phi\) on \(X\), \[ \int_X \Phi(x)\,d\mu(x)=\int_Y\left(\int_{f^{-1}(y)} \Phi(x)\,d\mu_y(x)\right)\,d(f_\*\mu)(y). \] This formula shows how pushforward supplies the weights on \(Y\), while the fiber measures supply the internal distribution on each preimage.
7.3.1 Fiber measures in countable or smooth settings
In settings with countable fibers or smooth structures allowing regular conditional laws, the conditional measures \(\mu_y\) can sometimes be described more concretely. For instance, in countable-to-one maps, conditional measures distribute mass across the finitely or countably many preimages in a way compatible with the original density.
8 Typical examples and canonical constructions
Pushforwards appear throughout measure theory and probability in concrete, computable forms.
8.1 Dirac measures and images of points
For a point \(x\in X\), the Dirac measure \(\delta_x\) satisfies \(\delta_x(A)=1\) if \(x\in A\) and \(0\) otherwise. Its pushforward is the Dirac measure at the image: \[ f_\*\delta_x=\delta_{f(x)}. \] This example illustrates that pushforward generalizes “mapping points into the target” at the level of measures.
8.2 Counting measures under mappings
If \(\mu\) is counting measure on \(X\), then \(f_\*\mu\) assigns to each \(B\subseteq Y\) the number of points in \(X\) mapping into \(B\), possibly infinite. In finite-to-one contexts, this becomes a weighted counting measure on \(Y\) determined by fiber sizes.
8.3 Lebesgue measure under affine maps
| Let \(f(x)=Ax+b\) be an affine map on \(\mathbb{R}^n\). Under suitable linear-algebraic assumptions (such as \(A\) invertible), the pushforward of Lebesgue measure is another Lebesgue-type measure scaled by \( | \det A | \). This is a prototypical instance of the Jacobian mechanism. |
|---|
8.3.1 Scaling and translation effects
Translation \(x\mapsto x+b\) does not change Lebesgue measure, while linear scaling contributes a determinant factor. Hence \(f_\*\lambda\) reflects only the volume distortion created by the linear part \(A\).
8.4 Pushforwards under projections in product spaces
In product spaces, projections map \((x,y)\mapsto x\) (or to \(y\)). Pushing forward under such a projection yields marginal measures.
8.4.1 Marginalization as a pushforward
Given a measure \(\mu\) on \(X\times Y\), the pushforward along the projection \(\pi_X\) produces the \(X\)-marginal: \[ (\pi_X)_*\mu(A)=\mu(A\times Y). \] This is the measure-theoretic expression of “marginal distribution” in probability.
9 Applications in analysis and probability
Pushforwards unify change-of-variables, distribution transformations, and stability questions under limits.
9.1 Distribution transformations in probability
If random variables are defined on a common probability space, their joint law can be pushed forward under coordinate transformations to obtain new laws. For example, applying a measurable function to one coordinate yields the distribution of that transformed coordinate, computed via the pushforward of the original law.
9.2 Weak convergence and continuity of pushforward
Pushforward interacts with weak convergence through continuity properties. If measures \(\mu_n\) converge weakly to \(\mu\) on \(X\), and \(f\) is continuous (with appropriate tightness or boundedness conditions), then \(f_\*\mu_n\) converges weakly to \(f_\*\mu\) on \(Y\).
9.2.1 Convergence in distribution via pushforwards
In probabilistic language, this means that if random elements \(Z_n\) converge in distribution to \(Z\), then the transformed elements \(f(Z_n)\) converge in distribution to \(f(Z)\), provided \(f\) is continuous at the relevant points. The pushforward provides the formal measure statement behind this principle.
9.3 Support and image sets
The support of a pushforward measure is closely tied to the image of the original measure through \(f\). Informally, the pushforward places mass only on points in \(Y\) that are approached by images of sets with positive \(\mu\)-mass.
9.3.1 Where the pushed measure “lives”
If \(\mu\) is concentrated on a set \(E\subseteq X\), then \(f_\*\mu\) is concentrated on \(f(E)\) in the measurable sense. This is often used to restrict attention to the effective range of a transformed random variable or an induced law in analysis.
9.4 Transport-like interpretations (lightweight conceptual overview)
Conceptually, pushforward can be viewed as a mechanism for reorganizing “mass” without changing total amount: it transfers the weight of each part of the domain to the part of the codomain determined by the map.
9.4.1 From measures to induced laws of transformed variables
In applications, one typically starts with a base measure describing uncertainty or distribution on \(X\), then applies a transformation to obtain a law on \(Y\). Pushforward is the formal bridge between the original description and the induced distribution of the transformed quantity.