University Physics V · Mathematical Foundations of Quantum Physics · 1.4
Orthogonality, normalisation and basis expansion
Quantum mechanics answers every question by projecting a state onto a basis. This topic is where you earn the right to do that: normalise so the numbers are probabilities, orthogonalise so they never double-count, and audit completeness so they sum to one.
Build the model
Connect the measurement to the mechanism.
The Born rule promises that |⟨n|ψ⟩|² is a probability, and that promise is bought entirely by the structure of the basis — this topic is the purchase. Normalisation, ⟨ψ|ψ⟩ = 1, spends the free scale in a ray so that squared moduli read as fractions of certainty. Orthogonality, ⟨m|n⟩ = δₘₙ, is what lets one inner product cₙ = ⟨n|ψ⟩ pull out one coefficient: apply ⟨m| to the expansion and the Kronecker δ kills every term but its own, where a non-orthogonal set would force a Gram-matrix solve.
Parseval's identity, Σ|cₙ|² = ⟨ψ|ψ⟩ = 1, then converts normalisation into the statement that the outcome probabilities exhaust certainty. Gram–Schmidt is the constructive guarantee that such bases exist: any independent set — say the degenerate eigenvectors a Hermitian operator returns in no particular arrangement — is orthonormalised by subtracting projections and rescaling, without ever leaving the subspace it spans. The cost is the one thing the algebra cannot check: completeness.
An orthonormal set that spans too little still yields impeccable coefficients; it simply returns Σ|cₙ|² < 1 and loses the rest of the probability without protest, which is why every truncated expansion, analytic or numerical, ends with a Parseval audit.
- Simple definition
- An orthonormal basis satisfies ⟨m|n⟩ = δₘₙ, so any state expands as |ψ⟩ = Σ cₙ|n⟩ with cₙ = ⟨n|ψ⟩ read off by projection, and normalising to ⟨ψ|ψ⟩ = 1 makes each |cₙ|² a measurement probability.
- Example
- For |ψ⟩ = (|1⟩ + i√3|2⟩)/2 on an orthonormal (|1⟩, |2⟩), projection gives c₁ = 1/2 and c₂ = i√3/2, so P₁ = 1/4 and P₂ = 3/4, summing to 1; the factor i moves no probability in this basis, though it would in a rotated one.
The Kronecker delta is a selection rule: it is what turns a spanning set into a coordinate system, where each direction can be read without reference to the others.
δₘₙ = 1 for m = n and 0 otherwise; continuum bases replace it with δ(x − x′) and live outside L²
Fixes the scale of the ray; the freedom left over is exactly one global phase e(iα), which no measurement sees.
in 1D position space the normalised ψ(x) carries units m⁻¹⁄², so |ψ|² dx is a pure number
Turns "which state is this?" into a list of complex numbers in a chosen basis — the coordinates quantum mechanics computes with.
cₙ ∈ ℂ; the sum runs over the whole basis — the resolution of the identity Σ|n⟩⟨n| = 1 does the work
The audit line: a kept set returning Σ|cₙ|² = 0.94 has silently discarded 6% of the probability.
cross terms cₘ*cₙ⟨m|n⟩ die by orthogonality; equality with 1 needs both normalisation and a complete set
Manufactures an orthonormal basis from any independent set — run inside a degenerate eigenspace it never leaves the eigenspace.
subtract the components already counted, then rescale; ‖wₙ‖ = 0 flags a dependent vector, which drops
Shows what orthonormality buys: G = 1 turns a linear solve into a single inner product per coefficient.
G is Hermitian and positive-definite for an independent set; det G → 0 warns of near-dependence
Normalisation fixes the scale — and only the scale
A physical state is a ray: |ψ⟩ and z|ψ⟩ describe the same system for any nonzero complex z. Demanding ⟨ψ|ψ⟩ = 1 spends that freedom down to a single unobservable global phase, and it is what licenses reading squared moduli as probabilities. The mechanics is one inner product: |φ⟩ = |1⟩ + 2i|2⟩ has ⟨φ|φ⟩ = 1⋅1 + (−2i)(2i) = 1 + 4 = 5 — conjugate in the first slot, or the crossed i's give −4 and an impossible negative norm — so the normalised state is (|1⟩ + 2i|2⟩)/√5. On L² the same step fixes dimensions: for ψ(x) = A sin(πx/L) on [0, L], ∫|A|² sin²(πx/L) dx = |A|²L/2 = 1 gives A = √(2/L), and ψ carries units of m⁻¹⁄² so that |ψ|² dx is a pure number. Choosing A real and positive is convention, not physics: √(2/L)⋅e(iα) is exactly as normalised.
Orthogonality lets one inner product read one coefficient
Expand |ψ⟩ = Σₙ cₙ|n⟩ and apply ⟨m|: linearity gives ⟨m|ψ⟩ = Σₙ cₙ⟨m|n⟩, and if ⟨m|n⟩ = δₘₙ the sum collapses to the single term cₘ. That collapse is the whole convenience of an orthonormal basis — each coefficient is extracted independently, by one integral or one dot product, with no reference to the others. Remove orthogonality and the extraction entangles: for an independent but non-orthogonal set (|vₙ⟩), the same move gives ⟨vₘ|ψ⟩ = Σₙ Gₘₙcₙ with Gₘₙ = ⟨vₘ|vₙ⟩, a linear system to solve. Try |v₁⟩ = |1⟩, |v₂⟩ = (|1⟩ + |2⟩)/√2 and |ψ⟩ = |1⟩: the raw projections are 1 and 1/√2, but solving with G = [[1, 1/√2], [1/√2, 1]] returns c = (1, 0) — the second vector contributes nothing, even though ψ has a healthy overlap with it. Overlap and coefficient only coincide when the basis is orthonormal.
Parseval turns the norm into total probability
Compute ⟨ψ|ψ⟩ from the expansion and both slots bring a sum: ⟨ψ|ψ⟩ = Σₘ Σₙ cₘ* cₙ ⟨m|n⟩, an N²-term double sum. Orthonormality deletes every off-diagonal term and leaves Parseval's identity, Σₙ|cₙ|² = ⟨ψ|ψ⟩ = 1. This is the bookkeeping behind the Born rule: for a non-degenerate eigenvalue aₙ the numbers P(aₙ) = |cₙ|² are non-negative and now provably sum to one, so they form a probability distribution over measurement outcomes, with ⟨A⟩ = Σₙ aₙ|cₙ|² following at once. Read in reverse it is a diagnostic: probabilities that refuse to sum to one mean the basis is not orthonormal, the state is not normalised, or terms are missing. And it degrades gracefully — for any orthonormal subset, Σ|cₙ|² ≤ 1 (Bessel's inequality), the deficit equalling the probability of finding the system outside the spanned subspace. That inequality becoming equality for every state is exactly what completeness means.
Gram–Schmidt subtracts what is already counted
Hermitian operators hand you orthogonality only across distinct eigenvalues; inside a degenerate eigenspace, the eigenvectors that fall out of solving (A − a)v = 0 by hand are merely independent. Gram–Schmidt repairs them. Normalise the first: |e₁⟩ = |v₁⟩/‖v₁‖. From the second, subtract its component along e₁: |w₂⟩ = |v₂⟩ − ⟨e₁|v₂⟩|e₁⟩, so ⟨e₁|w₂⟩ = ⟨e₁|v₂⟩ − ⟨e₁|v₂⟩ = 0 by construction; then normalise. Each later vector is stripped of every direction already banked before it is admitted. A zero remainder ‖wₙ‖ = 0 is information, not failure: vₙ was dependent on its predecessors and drops. Because every |eₖ⟩ is a combination of the degenerate eigenvectors, the new basis never leaves the eigenspace — orthonormalisation costs nothing in eigenvalue terms. Run on 1, x, x² with the inner product ∫₋₁¹ f*g dx, the same algorithm manufactures the Legendre polynomials up to normalisation; the ordering matters, and each ordering yields a different, equally legitimate basis.
Completeness is an assumption — attach a number to it
No amount of orthonormality certifies that a set spans the space. Coefficients against an incomplete orthonormal set are each individually correct — still ⟨n|ψ⟩ — but Σ cₙ|n⟩ rebuilds only the projection of ψ onto the span, and the missing probability 1 − Σ|cₙ|² vanishes without an error message. In practice completeness is imported as a theorem — the spectral theorem in finite dimensions, Sturm–Liouville theory for eigenfunctions on an interval — and what you actually check is the truncation. How fast that audit converges is set by the boundary conditions: the flat state ψ(x) = 1/√L on [0, L] is square-integrable but non-zero at walls where every sine eigenfunction vanishes, so its odd-n coefficients cₙ = 2√2/(nπ) decay only as 1/n. The ground state captures |c₁|² = 8/π² = 0.811, adding n = 3 reaches 0.901, and Σ 8/(n²π²) over odd n crawls to 1 with a Gibbs overshoot at each wall — convergence in the mean-square norm, never pointwise. A smooth state that respects the walls converges orders of magnitude faster, and parity halves the work either way: a state even about the well centre has no even-n coefficients at all.
In NumPy the whole construction is one QR factorisation
Stack the vectors as columns of A; then Q, R = np.linalg.qr(A) returns in Q an orthonormal basis for the same span, with R recording the projections Gram–Schmidt would have subtracted. Use it instead of coding the textbook loop: classical Gram–Schmidt in floating point loses orthogonality in proportion to κ(A)²ε for nearly dependent columns, the modified variant to κ(A)ε, while Householder QR keeps ‖Q†Q − I‖ at machine level regardless. Two habits keep grid work honest. First, the inner product must carry the measure: ⟨f|g⟩ ≈ Σⱼ f*(xⱼ)g(xⱼ)Δx, so vectors orthonormal under NumPy's plain dot product represent functions mis-scaled by √Δx — build the weight in before comparing with analytic states. Second, audit rather than trust: np.max(np.abs(Q.conj().T @ Q − np.eye(k))) should sit near 10⁻¹⁵, and Parseval on a known state should return 1 to the same tolerance, before any spectrum computed in that basis is believed.
Change one variable at a time
Make the relationship visible.
With φ at 35°, drag θ down from 90° and the bar climbs past the Parseval line because both projections count the direction the two basis vectors share; θ = 20° with φ = 10° pushes the naive sum to 1.94, θ = 20° with φ = 80° drops it to 0.28, and θ = 90° pins it to 1 for every φ.
NAIVE c₁ = ⟨v₁|ψ⟩0.82
NAIVE c₂ = ⟨v₂|ψ⟩0.57
NAIVE Σ|cₖ|²1.000
GRAM–SCHMIDT Σ|cₖ|²1.000
Live interpretationNAIVE c₁ = ⟨v₁|ψ⟩: 0.82. NAIVE c₂ = ⟨v₂|ψ⟩: 0.57. NAIVE Σ|cₖ|²: 1.000. GRAM–SCHMIDT Σ|cₖ|²: 1.000
Catch the common trap
Explain before calculating.
The state |ψ⟩ = N(2|1⟩ − 2i|2⟩ + |3⟩) is written on an orthonormal basis (|1⟩, |2⟩, |3⟩), the non-degenerate eigenbasis of some observable. With N chosen to normalise it, what is the probability that measuring that observable returns the eigenvalue belonging to |2⟩?
Choose an answer to test the model.
Practice & worked examples
Reason from the model, then test the result.
EasyA two-state system has orthonormal basis (|1⟩, |2⟩) and is prepared in the unnormalised state |φ⟩ = 3|1⟩ − 4i|2⟩. Normalise it and find the probability of each measurement outcome.
- The norm is a sum of conjugate-squares: ⟨φ|φ⟩ = |3|² + |−4i|² = 9 + 16 = 25. Squaring −4i without conjugating would return 9 − 16 = −7, which no norm can be — the first flag that a conjugate was dropped.
- Divide by √25 = 5: |ψ⟩ = (3|1⟩ − 4i|2⟩)/5, so c₁ = 3/5 and c₂ = −4i/5.
- Probabilities are squared moduli: P₁ = 9/25 = 0.36 and P₂ = 16/25 = 0.64, and they sum to 1 — Parseval on a two-term basis.
- Any global phase is equally valid: e(iα)|ψ⟩ gives the same P₁ and P₂. Only the relative phase between c₁ and c₂ (the −i) can ever matter, and only in a rotated measurement basis.
Answer|ψ⟩ = (3|1⟩ − 4i|2⟩)/5; P₁ = 9/25 = 0.36, P₂ = 16/25 = 0.64, summing to 1.
MediumA Hermitian operator on ℂ³ has a doubly degenerate eigenvalue whose eigenspace is spanned by v₁ = (1, 1, 0)ᵀ and v₂ = (1, 0, 1)ᵀ — independent but not orthogonal, since ⟨v₁|v₂⟩ = 1. Run Gram–Schmidt to produce an orthonormal basis of the eigenspace.
- Normalise the first vector: ‖v₁‖² = 1 + 1 + 0 = 2, so e₁ = (1, 1, 0)ᵀ/√2.
- Project the second onto it: ⟨e₁|v₂⟩ = (1⋅1 + 1⋅0 + 0⋅1)/√2 = 1/√2. Subtract that component: w₂ = v₂ − (1/√2)e₁ = (1, 0, 1)ᵀ − ½(1, 1, 0)ᵀ = (½, −½, 1)ᵀ.
- Check orthogonality before normalising: ⟨e₁|w₂⟩ = (½ − ½ + 0)/√2 = 0, as the construction guarantees.
- Normalise the remainder: ‖w₂‖² = ¼ + ¼ + 1 = 3/2, so e₂ = (½, −½, 1)ᵀ/√(3/2) = (1, −1, 2)ᵀ/√6.
- Both e₁ and e₂ are linear combinations of v₁ and v₂, so both are still eigenvectors with the same eigenvalue: Gram–Schmidt reorganised the eigenspace without leaving it. Ordering v₂ first would give a different, equally valid pair.
Answere₁ = (1, 1, 0)ᵀ/√2 and e₂ = (1, −1, 2)ᵀ/√6 — orthonormal, and still inside the degenerate eigenspace.
HardA particle in an infinite well of width L is prepared in ψ(x) = √(30/L⁵) x(L − x). Expand it on the energy eigenbasis φₙ(x) = √(2/L) sin(nπx/L), find the probability of measuring E₁, and use Parseval's identity to audit a truncation that keeps only n = 1.
- Check normalisation first: ∫₀ᴸ x²(L − x)² dx = L⁵(1/3 − 1/2 + 1/5) = L⁵/30, so ⟨ψ|ψ⟩ = (30/L⁵)(L⁵/30) = 1.
- Project onto each eigenstate: cₙ = ⟨φₙ|ψ⟩ = (√60/L³) ∫₀ᴸ x(L − x) sin(nπx/L) dx = (√60/L³) · 2L³(1 − cos nπ)/(n³π³). Even n vanish — ψ is even about L/2 while the even-n eigenfunctions are odd about it — and odd n give cₙ = 8√15/(n³π³).
- The ground-state probability is P(E₁) = |c₁|² = 64⋅15/π⁶ = 960/π⁶ = 0.99855.
- Parseval audits the truncation: keeping n = 1 alone loses 1 − 0.99855 = 0.00145 of the probability, and the next allowed term carries |c₃|² = 960/(3⁶π⁶) = 0.00137 — about 95% of that deficit.
- Completeness closes the books exactly: Σ over odd n of 1/n⁶ = π⁶/960, so Σ|cₙ|² = (960/π⁶)(π⁶/960) = 1.
AnswerP(E₁) = 960/π⁶ ≈ 0.9986. Truncating to n = 1 silently loses 0.14% of the probability, and the n = 3 term carries about 95% of that deficit.