ISEGORIA / MATH ENCYCLOPEDIA
Push-forward and pullback: integrating through a map
How integrals travel along a map: pull a function or a form back and integrate upstairs, or push a measure forward and integrate downstairs. The area formula and its multiplicity, push-forward measures and their densities, and integration along the fibres of a map.
Before you begin: Multivariable calculus, Jacobians, and a first look at measures and differential forms
Predict, manipulate, then check your reasoning against the example and question. Graphs illustrate the mathematics; they do not replace a proof.
1. Pull back, then integrate
A map φ from the unit square onto a region lets you compute an integral downstairs by pulling it back upstairs: replace f by f∘φ and dx dy by |det Dφ| du dv. Here φ is the bilinear map through four corners you can drag. While φ is one-to-one every version of the integral agrees. Drag the gold corner φ(1,1) inside the triangle formed by the other three and φ folds over: part of the plane is hit twice, and the pulled-back integral counts it twice. That is the area formula, with the multiplicity #φ⁻¹(y) made visible.
Worked example. Send the square to the parallelogram with corners \(0,(2,0),(3,1),(1,1)\): \(\varphi(u,v)=(2u+v,\,v)\) and \(\det D\varphi=2\) everywhere, so \(\int_{[0,1]^2}2\,du\,dv=2\), the area of the parallelogram. With \(f=x\) the pulled-back integrand is \((2u+v)\cdot2\), which integrates to \(3\), and \(3/2\) is indeed the \(x\) coordinate of the centre \((3/2,\,1/2)\).
Watch out. Two different objects are being pulled back. The measure-theoretic change of variables uses \(|\det D\varphi|\) and counts preimages. The pullback of the area form \(f\,dx\wedge dy\) keeps the sign, so its integral counts preimages with orientation, and a folded layer cancels. They agree exactly when \(\varphi\) is one-to-one and keeps orientation. For a bilinear map \(\det D\varphi\) is affine in \((u,v)\), so the fold is a straight segment.
Why does the signed integral of det Dφ equal the shoelace area of the four corners, even when the map folds?
Because \(\int_D\varphi^*(dx\wedge dy)=\int_D d\,\varphi^*(x\,dy)=\oint_{\partial D}\varphi^*(x\,dy)\) by Stokes, and the last integral depends only on the boundary curve \(\varphi(\partial D)\). A bilinear map sends each edge of the square to a straight segment, and \(\oint x\,dy\) around a polygon is the shoelace sum.
2. Pushing a measure forward
A measurable map T carries a measure μ on its domain to a measure T#μ on its target, defined by pulling sets back: T#μ(B) = μ(T⁻¹(B)). Integrals then move the other way, ∫ g d(T#μ) = ∫ (g∘T) dμ, so pushing measures forward and pulling functions back are adjoint operations. Here T(x) = x + a sin(2πx)/(2π) on [0, 1]. Drag the window B and watch its preimage split into pieces once the graph folds; drop particles sampled from μ and watch them pile up where T is flat.
Worked example. With \(\mu\) uniform and \(a<1\), \(T^{\prime}(x)=1+a\cos2\pi x>0\), so the density of \(T_{\#}\mu\) at \(y=T(x)\) is \(1/(1+a\cos2\pi x)\). It is thinnest, \(1/(1+a)\), at \(y=0\) and \(y=1\) where \(T\) is steepest, and thickest, \(1/(1-a)\), at \(y=\tfrac12\) where \(T\) is flattest.
Watch out. The push-forward needs only \(T\) and \(\mu\) and always exists. A density for it exists when \(T^{\prime}\ne0\) almost everywhere, and even then it is a sum over every preimage, not just one. There is no general way to pull a measure back along a map: measures push forward, functions and forms pull back.
For a > 1, why does the density blow up like 1/√|y − y*| at a fold height y*, while the mass near y* stays finite?
Near a turning point \(x^*\), \(T(x)\approx y^*+\tfrac12T^{\prime\prime}(x^*)(x-x^*)^2\). The two preimages of \(y\) lie at distance \(\sqrt{2|y-y^*|/|T^{\prime\prime}|}\) from \(x^*\), where \(|T^{\prime}|\approx\sqrt{2|T^{\prime\prime}|\,|y-y^*|}\). Dividing \(\rho\) by that gives the inverse square root, and \(\int|y-y^*|^{-1/2}dy\) converges.
3. Integrating along the fibres
Push a density p on the plane forward by a function F: ℝ² → ℝ and you get the distribution of F(X) when X has density p. Its density at c comes from integrating p over the fibre F = c, weighted by 1/|∇F|, because the shell between the fibres at c and c + δ has width δ/|∇F|. This is the coarea formula. Marginals, convolutions, the Rayleigh law and the density of a product are all instances of it. Choose the map, then drag the level c and watch the shell.
Worked example. For the standard Gaussian and \(F=\sqrt{x^2+y^2}\): the fibre is a circle of length \(2\pi c\), \(|\nabla F|=1\), and \(p=e^{-c^2/2}/2\pi\) on it, so \(q(c)=c\,e^{-c^2/2}\), the Rayleigh density. It vanishes at \(c=0\) although \(p\) is largest there.
Watch out. The formula needs \(\nabla F\ne0\) on almost all of each fibre. At a critical value, such as \(c=0\) for \(F=xy\), the fibre runs through a critical point and \(q\) can be unbounded while its integral stays finite. The bars on the right use no calculus at all: they count mass cell by cell on the plane, so they are an independent check on the curve.
Why is the density of X + Y the convolution ∫ p(t, c − t) dt, and not that integral divided by √2?
Parametrising the line \(x+y=c\) by \(x=t\) gives arc length \(d\ell=\sqrt2\,dt\), and \(|\nabla(x+y)|=\sqrt2\). The two factors of \(\sqrt2\) cancel.