University Physics V · Mathematical Foundations of Quantum Physics · 1.2
Probability Distributions, Moments & Expectation Values
Every number quantum mechanics predicts is an integral against a weight over outcomes. This topic drills that machinery on classical weights — a die, a decay time, a box density — so that checking normalisation and units, computing moments and quoting a spread are reflexes by the time Born's rule swaps in |ψ|².
Build the model
Connect the measurement to the mechanism.
Quantum mechanics never hands you an outcome; it hands you a weight — a rule assigning each possible result its share of an infinitely repeated experiment — and everything the theory exports is an integral against that weight. This topic builds the export machinery while the weight is still classical, so nothing quantum obscures it. A discrete distribution is a list Pₙ summing to 1; a continuous density ρ(x) is subtler — a probability per unit x, carrying inverse units and paying out actual probabilities only over intervals.
Every prediction is then a moment: the mean ⟨x⟩ = ∫x ρ dx locates the distribution, the variance σ² = ⟨x²⟩ − ⟨x⟩² prices its spread, and ⟨f(x)⟩ weights any function outcome by outcome — which is generally not f(⟨x⟩). Two costs come with the machinery. First, it is strictly ensemble talk: the mean is an average over identically prepared systems and need not be a possible outcome of any single run, so intuition trained on "the expected value" misleads.
Second, existence is not guaranteed: a Lorentzian line shape is a perfectly normalised density with no variance at all. In Unit 2 Born's rule sets ρ = |ψ|², and Unit 3 rewrites every one of these integrals as ⟨ψ|Â|ψ⟩ — the same machinery, loaded with a new weight.
- Simple definition
- A probability distribution assigns each outcome its long-run share over identical trials — dimensionless Pₙ for discrete outcomes, a density ρ(x) with the inverse unit of x for continuous ones — and moments ⟨xᵏ⟩ = ∫xᵏρ dx compress that weight into location, spread and shape.
- Example
- A fair die has Pₙ = 1/6 for n = 1…6: normalisation 6 × 1/6 = 1, mean ⟨n⟩ = 3.5 — a value no roll can show — ⟨n²⟩ = 91/6 = 15.17, and σ = √(15.17 − 12.25) = 1.71.
Run this check first: it fixes any free constant in the weight and reads off ρ's units before a single moment is computed.
Pₙ is a pure number; ρ(x) carries the inverse unit of x (m⁻¹ for position, s⁻¹ for a lifetime), so ρ(x) dx is dimensionless
Locates the distribution's balance point — which can sit where ρ = 0, so never read it as the likely value.
⟨x⟩ shares x's unit; xₙ are the outcomes, Pₙ their weights — an ensemble average, not any single run's result
The template every quantum prediction fills: kinetic energy, potential energy and every ⟨Â⟩ are functions averaged against a weight.
f is weighted outcome by outcome; ⟨f(x)⟩ = f(⟨x⟩) is guaranteed only for linear f, and ⟨x²⟩ ≥ ⟨x⟩² always
σ prices the spread of single outcomes about the mean — the Δx and Δp bounded by every uncertainty relation are exactly this σ.
σ carries x's unit; expanding the square collapses the cross term, since ⟨x⟩ is a constant inside the average
Units 2 and 3 load this weight into today's machinery unchanged: ⟨x⟩ = ∫x|ψ|² dx is the same integral, later rewritten as ⟨ψ|x̂|ψ⟩.
in 1D ψ carries units of m⁻¹⁄², so |ψ|² is a genuine density; cₙ = ⟨n|ψ⟩ are basis-expansion coefficients
Spectral lines and resonances take this shape — characterise them by Γ, because the second moment does not exist to be quoted.
the Lorentzian: x₀ centres it, Γ is the full width at half maximum; ∫ρ dx = 1 exactly, yet σ² diverges
Start discrete: a complete list of outcomes and weights
The cleanest case first: an experiment with countably many outcomes x₁, x₂, … and a number Pₙ attached to each, meaning the fraction of an unbounded run of identical trials that returns xₙ. Two axioms do all the work: every Pₙ ≥ 0, and Σ Pₙ = 1. The second is a completeness statement — the outcome list misses nothing — and it is the line to check first whenever a distribution is handed to you, because a list summing to 0.98 is silently losing one trial in fifty. The mean weights each outcome by its share, ⟨x⟩ = Σ xₙPₙ, and already carries this topic's central warning: a fair die has ⟨n⟩ = (1+2+3+4+5+6)/6 = 3.5, and no face shows 3.5. The mean is a property of the ensemble, never a prediction for one roll. Quantum mechanics will reuse this case verbatim: measuring an observable with a discrete, non-degenerate spectrum returns eigenvalue aₙ with probability Pₙ = |cₙ|², and the requirement Σ|cₙ|² = 1 is exactly this normalisation, resurfacing in Topic 1.4 as Parseval's identity.
A density is probability per unit length, not probability
When the outcome is continuous — a position, an arrival time — the probability of any exact value is zero: uncountably many candidates cannot each carry a finite share. The workable object is a density: ρ(x) dx is the probability of landing between x and x + dx, and only integrals of ρ are probabilities, P(a < x < b) = ∫ₐᵇ ρ dx. Two consequences trip students every year. First, ρ carries units — m⁻¹ for a position density, s⁻¹ for a lifetime density — because it must integrate against dx to a pure number. Second, the height of ρ is bounded by nothing: ρ(x) = 3x² on 0 ≤ x ≤ 1 m is perfectly normalised, ∫₀¹ 3x² dx = 1, yet reaches ρ = 3 m⁻¹ at the right edge. Only the area obeys the ≤ 1 rule; the height does whatever it must to hold the area. The Dirac delta bridges the two languages: a discrete outcome at x₀ with weight P is the density term P δ(x − x₀) — the form in which mixed spectra, bound levels plus a continuum, will be written later in the course.
Moments: every export is an integral against the weight
Every number a distribution can export is one integral: ⟨f(x)⟩ = ∫ f(x) ρ(x) dx, the value of f at each outcome weighted by that outcome's share. The powers f = xᵏ are the moments. For ρ = 3x² on [0, 1 m]: ⟨x⟩ = ∫₀¹ 3x³ dx = 3/4 m. Averaging does not commute with nonlinear functions: ⟨x²⟩ = ∫₀¹ 3x⁴ dx = 3/5 = 0.600 m², while ⟨x⟩² = 0.5625 m² — and the gap is no accident, since ⟨x²⟩ − ⟨x⟩² = ⟨(x − ⟨x⟩)²⟩ ≥ 0, with equality only for a δ-sharp weight. Keep mean, median and mode distinct: here the mode is 1 m where ρ peaks, the median m solves m³ = 1/2 giving 0.794 m, and the mean is 0.750 m. Skew pulls the three apart — here, piled against the right wall, in the order mean < median < mode — and a symmetric single-peaked density brings them back together. Symmetry is also the working shortcut: when ρ is even about a point, the integrand (x − x₀)ρ is odd and the mean sits at x₀ by cancellation — the one-line argument that will give ⟨x⟩ = L/2 for every infinite-well eigenstate before any integral is attempted.
Variance: two forms of one number, and when each wins
The definitional form σ² = ⟨(x − ⟨x⟩)²⟩ says what variance is: the mean squared deviation from the mean. Expanding the square gives ⟨x²⟩ − 2⟨x⟩⟨x⟩ + ⟨x⟩² = ⟨x²⟩ − ⟨x⟩², the computational form, quickest by hand: for ρ = 3x² on [0, 1 m], σ² = 3/5 − 9/16 = 3/80 = 0.0375 m², so σ = 0.194 m. Numerically the shortcut can be a trap: a narrow density parked at a large mean — a 1 nm spread on a 10 μm mean — makes ⟨x²⟩ and ⟨x⟩² agree to eight significant figures, and a float64 subtraction hands back only the eight that remain; numpy.var centres the data before squaring for exactly this reason. Quote σ, not σ², as the physical width: it shares x's unit. And keep it distinct from an instrument's error bar — this σ is intrinsic to the distribution and survives perfect apparatus. The Δx in Δx Δp ≥ ħ/2 is precisely this standard deviation, computed later from the weight |ψ|².
Heavy tails: normalisation does not buy moments
The Lorentzian ρ(x) = (Γ/2π)/[(x − x₀)² + (Γ/2)²] integrates to exactly 1, yet its tails fall only as 1/x²: the variance integrand x²ρ tends to a constant, so σ² diverges, and even ⟨x⟩ exists only as a symmetric principal value. This is no invented pathology — it is the natural line shape of an exponentially decaying state, and it returns as the Breit–Wigner resonance in the particle-physics units, which is why resonances are quoted by their full width at half maximum Γ and never by a standard deviation. The practical symptom is that averaging fails: the mean of N independent Cauchy draws is distributed exactly like a single draw, so a simulation that tries to estimate ⟨x⟩ wanders forever instead of settling. Before quoting any moment, check the tail: ⟨xᵏ⟩ exists only if ρ falls faster than |x|⁻ᵏ⁻¹ as |x| → ∞.
The handoff: ensemble statistics and the Born weight
Everything above is classical probability, and that is the point: quantum mechanics changes where the weight comes from, not what you do with it. Born's rule — postulated in Unit 2 and applied to position in Unit 3 — will declare ρ(x) = |ψ(x)|² for position and Pₙ = |cₙ|² for a discrete spectrum; every formula on this page then applies verbatim, merely rewritten there in operator dress, ⟨x⟩ = ⟨ψ|x̂|ψ⟩. What the theory predicts is therefore ensemble statistics: prepare N identical systems, measure each once, and the histogram converges on ρ while the sample mean approaches ⟨x⟩ with statistical scatter dying as σ/√N. No formula here addresses a single run — not a week-one gap to be repaired later, but the permanent shape of the theory's output. One NumPy line makes it concrete: rng.choice([1, 2, 3], p=[0.2, 0.5, 0.3], size=10000) has sample means scattering about the exact ⟨a⟩ = 2.1 with spread σ/√N = 0.7/100 = 0.007 — the 1/√N crawl every Monte Carlo lab in this course will exhibit.
Change one variable at a time
Make the relationship visible.
Keep w = 0.50 and drag s from 0 to 2 nm: the dashed mean never leaves x = 0 even as the density there collapses toward zero, while σ grows past 2 nm — moments measure balance and spread, not likelihood. Then push w toward 0.95 and watch the mean slide into the heavier hump.
MEAN ⟨x⟩0.00 nm
SECOND MOMENT ⟨x²⟩1.56 nm²
STD DEV σ1.25 nm
DENSITY AT MEAN0.003 nm⁻¹
Live interpretationMEAN ⟨x⟩: 0.00 nm. SECOND MOMENT ⟨x²⟩: 1.56 nm². STD DEV σ: 1.25 nm. DENSITY AT MEAN: 0.003 nm⁻¹
Catch the common trap
Explain before calculating.
The position of a particle is distributed with density ρ(x) = 3x² for 0 ≤ x ≤ 1 m and zero elsewhere. Which statement about measurements on this ensemble is correct?
Choose an answer to test the model.
Practice & worked examples
Reason from the model, then test the result.
EasyA measurement of an observable  on an ensemble of identically prepared systems returns eigenvalue a = 1, 2 or 3 (in units of ħω) with probabilities P₁ = 0.2, P₂ = 0.5, P₃ = 0.3. Confirm the distribution is normalised, compute ⟨a⟩, ⟨a²⟩ and σₐ, and state whether any single run can return ⟨a⟩.
- Normalisation: 0.2 + 0.5 + 0.3 = 1.00 — the outcome list is complete and no weight leaks.
- ⟨a⟩ = (1)(0.2) + (2)(0.5) + (3)(0.3) = 0.2 + 1.0 + 0.9 = 2.1 ħω.
- ⟨a²⟩ = (1²)(0.2) + (2²)(0.5) + (3²)(0.3) = 0.2 + 2.0 + 2.7 = 4.9 (ħω)² — square the outcomes, never the mean.
- σₐ² = ⟨a²⟩ − ⟨a⟩² = 4.9 − 4.41 = 0.49, so σₐ = 0.70 ħω.
- A single run returns an eigenvalue: 1, 2 or 3. The value 2.1 is unattainable in any one measurement — it emerges only as the N → ∞ average, approached at the rate σₐ/√N.
Answer⟨a⟩ = 2.1 ħω, ⟨a²⟩ = 4.9 (ħω)², σₐ = 0.70 ħω; no single run can return 2.1 — it is a property of the ensemble.
MediumUnit 5 will derive ρ(x) = C sin²(πx/L) on 0 ≤ x ≤ L as the position density of the infinite well's ground state. Treat it here as a given weight with L = 1.00 nm: find C, then ⟨x⟩, ⟨x²⟩ and the standard deviation σ.
- Normalise: sin² = (1 − cos(2πx/L))/2, and the cosine integrates to zero over the full interval, so ∫₀ᴸ sin²(πx/L) dx = L/2. Hence C = 2/L, with units m⁻¹ as a 1D density must have.
- ⟨x⟩ = L/2 = 0.500 nm by symmetry: ρ(L/2 + u) = ρ(L/2 − u), so the integrand (x − L/2)ρ is odd about the midpoint and integrates to zero — no integral needed.
- ⟨x²⟩ = (2/L)∫₀ᴸ x² sin²(πx/L) dx = (1/L)∫₀ᴸ x²(1 − cos(2πx/L)) dx. The first piece gives L²/3; integrating the cosine piece by parts twice gives ∫₀ᴸ x² cos(2πx/L) dx = L³/(2π²).
- So ⟨x²⟩ = L²(1/3 − 1/(2π²)) = L²(0.3333 − 0.0507) = 0.2827 L² = 0.283 nm².
- σ² = ⟨x²⟩ − ⟨x⟩² = L²(1/3 − 1/(2π²) − 1/4) = L²(1/12 − 1/(2π²)) = 0.0327 L², so σ = 0.181 L = 0.181 nm — the ground state parks about 18% of the box width either side of centre.
AnswerC = 2/L = 2.00 nm⁻¹; ⟨x⟩ = 0.500 nm; ⟨x²⟩ = 0.283 nm²; σ = 0.181 nm.
HardA decay time is distributed as ρ(t) = (1/τ) e(−t/τ) for t ≥ 0. Iodine-131 has half-life t₁/₂ = 8.02 d. Show that the median of ρ is t₁/₂ and find τ; then compute ⟨t⟩, ⟨t²⟩ and σₜ, and the fraction of nuclei that outlive 2⟨t⟩.
- Check the weight: ∫₀^∞ (1/τ)e(−t/τ) dt = 1 for any τ > 0, and ρ carries units of d⁻¹, so integrals against dt are pure probabilities. The cumulative is P(t < T) = 1 − e(−T/τ).
- Median m: 1 − e(−m/τ) = ½ gives m = τ ln 2. Half the ensemble has decayed by m, so m is precisely the half-life: τ = t₁/₂ / ln 2 = 8.02/0.6931 = 11.57 d.
- Moments via the Gamma integral ∫₀^∞ tᵏ e(−t/τ) dt = k! τᵏ⁺¹: ⟨t⟩ = (1/τ)(1!)τ² = τ = 11.57 d, and ⟨t²⟩ = (1/τ)(2!)τ³ = 2τ² = 267.7 d².
- σₜ² = 2τ² − τ² = τ², so σₜ = τ = 11.57 d: the exponential's spread equals its mean — 100% relative width — and the mean exceeds the median by the factor 1/ln 2 = 1.44, the long tail dragging the balance point right.
- P(t > 2⟨t⟩) = e(−2⟨t⟩/τ) = e⁻² = 0.135: two mean lives in, 13.5% of the nuclei still survive — unsurprising for this skewed weight, and exactly what symmetric intuition gets wrong.
Answerτ = ⟨t⟩ = σₜ = 11.6 d; ⟨t²⟩ = 2τ² = 268 d²; median = t₁/₂ = 8.02 d; P(t > 2⟨t⟩) = e⁻² = 13.5%.