Skip to main content
University Physics IV

University Physics IV · Frontiers of Modern Physics · 15.9

Reading a Frontier Claim

You will meet results that are not yet knowledge. This lesson supplies the arithmetic and the etiquette for judging them: what an exclusion curve actually forbids, why the field waits for five σ, how many searches you really ran, and which questions are still honestly open.

01

Build the model

Connect the measurement to the mechanism.

The output of a frontier experiment is almost never a discovery; it is a bound, and a bound is a real measurement. Sensitivity is arithmetic: an excess of s events over an expected background b is significant to roughly Z = s/√b, so significance grows only as the square root of running time, and every threshold a field adopts is a statement about how often it is willing to be fooled. Five sigma means a one-sided tail probability of 2.9 × 10⁻⁷ — but that is P(data this extreme | background only), never the probability that the signal is real, and it is computed inside an assumed background model, so a wrong model yields a confident wrong number.

Two corrections do most of the practical work. The look-elsewhere effect says that if you searched N independent places, the chance that the best of them fluctuates is about N times the local chance, which is why 3.5σ bumps in wide scans routinely evaporate. And the systematic budget imposes a ceiling: once the background uncertainty σb rivals √b, Z saturates at s/σb and further exposure buys nothing.

What actually converts a claim into knowledge is none of this — it is an independent group, with different systematics, finding the same thing. Everything before that is a number with a caveat attached, and physics keeps a standing list of caveats that have never resolved.

Simple definition
A frontier claim is a statement about data together with a stated probability of having been fooled by background alone; until an independent experiment with different systematics reproduces it, that probability is the entire content of the claim.
Example
A search expecting b = 900 background events needs s = 5√900 = 150 excess events for a local 5σ, a one-sided p of 2.9 × 10⁻⁷; if the scan covered 250 independent mass windows, the global p is 250 × 2.9 × 10⁻⁷ = 7.2 × 10⁻⁵, or 3.8σ.
Significance of a counting excessZ ≈ s / √b

b = 900 needs s = 150 for 5σ. With s and b both ∝ exposure, Z grows only as √exposure.

s = excess events above the expected background b; both are pure counts, valid for s ≪ b and b ≫ 1

The five-σ convention, one-sidedp(Z ≥ 3) = 1.3 × 10⁻³ · p(Z ≥ 5) = 2.9 × 10⁻⁷

1 in 740 is called evidence; 1 in 3.5 million is called an observation. Neither is a probability that the signal is real.

p is P(data this extreme | background only) — it is not P(background | data)

Look-elsewhere effect (trials factor)pglobal ≈ 1 − (1 − plocal)N ≈ N · plocal

With N = 250, a 5σ local becomes 3.8σ global, and a 3.5σ local becomes 1.6σ.

N = number of statistically independent resolution-width windows searched; needs N⋅plocal ≪ 1

Where the systematic budget caps a searchZ = s / √(b + σb²) → s₀/(f b₀) as exposure grows

A 3% systematic on b₀ = 100 events per unit caps Z at s₀/3 forever. More data cannot lift a ceiling.

σb = f⋅b is the background uncertainty in events; s = s₀t and b = b₀t for exposure t

Poisson upper limit with zero events seenμᵤₚ = −ln(1 − CL) · 2.30 at 90% CL · 3.00 at 95% CL

Seeing nothing is a measurement: every rate above μᵤₚ/(εL) is excluded at that confidence level.

μᵤₚ is in events; divide by efficiency × exposure to get a rate or cross-section limit

Tension between two measurementsT = |x₁ − x₂| / √(σ₁² + σ₂²)

H₀: |73.0 − 67.4| / √(0.5² + 1.0²) = 5.0σ — not a discovery, but a broken error budget.

Independent measurements of the same quantity; each σ in the same unit as x

01

A bound is a result, not a failure

Most frontier experiments never see anything, and that is the normal, publishable outcome. A search reports a curve: for each candidate mass, the largest signal rate compatible with what was seen, at a stated confidence level. Direct dark-matter searches have pushed the spin-independent WIMP–nucleon cross-section limit from about 10⁻⁴² cm² in the early 2000s to below 10⁻⁴⁷ cm² near 30 GeV — six orders of magnitude of parameter space closed without a single detection, and they are now approaching the irreducible background from solar and atmospheric neutrinos. Read the axes and the fine print together: that curve assumes a local halo density near 0.3 GeV cm⁻³ and a particular velocity distribution, so it excludes a model, not a particle. And "excluded at 90% CL" means the experiment would have seen more than this in 90% of repeats were the model true. It does not mean the model is 90% dead.

02

Five σ, and what a p-value is not

A p-value answers exactly one question: if only background were present, how often would a fluctuation this large or larger appear? For a one-sided Gaussian tail, 3σ gives p = 1.3 × 10⁻³, about 1 in 740, and 5σ gives 2.9 × 10⁻⁷, about 1 in 3.5 million. Particle physics calls 3σ "evidence" and reserves "observation" for 5σ, while a clinical trial is content with 0.05, and the gap is not fussiness. Three things drive it: thousands of searches run every year, so rare fluctuations are guaranteed somewhere; the prior on an exotic signal is low, so a modest p can leave a claim more likely wrong than right; and the historical record is full of 3σ effects that died. Note what the number does not cover. It is computed inside an assumed background model, and an effect the model omits shifts the central value while leaving the significance untouched — so no quantity of sigmas protects against it.

03

Count the places you looked

Do not ask how unlikely this bump is; ask how unlikely it is that the most extreme of all your searches produced one. If a scan covers N statistically independent windows, pglobal ≈ 1 − (1 − plocal)N ≈ N⋅plocal while that product stays small. Scan 100–2000 GeV with 20 GeV mass resolution and N ≈ 95: a local 3.7σ (p = 1.1 × 10⁻⁴) becomes a global p near 1.0 × 10⁻², which is 2.3σ, and 2.3σ is unremarkable. This is not hypothetical arithmetic. The 750 GeV diphoton excess of December 2015 was 3.9σ local and about 2.3σ global in ATLAS, 2.6σ local and under 1.2σ global in CMS; some five hundred theory papers followed, and the excess was gone in the next year's data. The informal version is worse — the cuts, binnings and variables you tried and did not record carry a trials factor nobody can reconstruct afterwards.

04

Blind the analysis, then budget the systematics

The uncountable trials factor has one honest remedy: fix everything before you look. In a blind analysis the cuts, background model and fitting procedure are frozen using simulation and sidebands, the signal region is opened once, and the result is published whatever it says. Variants add an unknown offset to the answer, or salt the data with fake signals — LIGO ran a blind hardware injection in 2010 that the collaboration analysed and wrote up before being told it was not real. Then the systematic budget: list every effect, size each one, and add independent contributions in quadrature. A budget of 1.2% energy scale, 1.0% background shape, 0.8% efficiency and 0.6% luminosity gives √(1.44 + 1.00 + 0.64 + 0.36) = 1.85%. Against a 0.40% statistical error the total is 1.90%. Quadruple the data and the total improves to 1.87% — a 1.7% gain for four times the running. That is where a measurement stops being statistics-limited, and no exposure repairs it.

05

Replication is what settles it

Nothing in the previous three sections converts a claim into knowledge. An independent experiment with different systematics does. OPERA's neutrinos arrived 60.7 ns early with ±6.9 ns statistical and ±7.4 ns systematic uncertainty — nominally about 6σ — and the cause was a loose fibre-optic connector together with a clock drift; ICARUS on the same beamline then measured a time consistent with c. BICEP2 announced r = 0.20 in 2014 at high significance, and a joint analysis using Planck's 353 GHz dust maps reassigned the signal to polarised galactic dust, leaving r < 0.12. The successes have the same shape. The Higgs was announced in July 2012 only when ATLAS and CMS each reached about 5σ independently; GW170817 became certain when an electromagnetic counterpart arrived 1.7 s later from an entirely different kind of instrument. Two recent claims of near-ambient superconductivity were withdrawn within months when other groups could not reproduce them.

06

The open list, written as open

Seven questions, each with the measured part separated from the unmeasured one. Neutrino oscillations fix only mass-squared differences — Δm²₂₁ ≈ 7.5 × 10⁻⁵ eV² and |Δm²₃₁| ≈ 2.5 × 10⁻³ eV², so the heaviest state is at least 0.050 eV — while the ordering and the absolute scale are unknown, bounded under about 0.5 eV directly and 0.12 eV by cosmology. The baryon-to-photon ratio is measured at 6 × 10⁻¹⁰; quark-sector CP violation is orders of magnitude too small to produce it. Ωₘ ≈ 0.31 against Ωb ≈ 0.049 leaves roughly 0.26 non-baryonic, with no candidate detected anywhere. ΩΛ ≈ 0.69 with w = −1.03 ± 0.03 is consistent with a constant, not proof of one. H₀ disagrees at about 5σ between early- and late-universe routes. The neutron EDM bound |dₙ| < 1.8 × 10⁻²⁶ e⋅cm forces the QCD angle below 10⁻¹⁰ for no known reason. And nothing tests gravity at 10¹⁹ GeV. Write each one as open; a forecast is not a result.

02

Change one variable at a time

Make the relationship visible.

Interactive model
20 events/unit
3.0 %
6 units

Hold the signal at 20 and drag the systematic from 0.5% to 5%: the curve stops climbing like √t and flattens onto its own ceiling below the 5σ line, and the exposure readout pins at 999 — a discovery no amount of running time can buy.

Interactive physics modelDiscovery significance against exposure for a counting search whose background rate is 100 events per unit. At t = 6 the excess stands at Z = 3.95σ, with 35% of the squared error already systematic; the dashed ceiling s₀/f = 6.7σ is the most any exposure can return.Z = s₀t / √( b₀t + (f b₀t)² ), b₀ = 100 per unitZ (σ)5σ discovery conventionceiling s₀/f = 6.7σ05100exposure t (0 → 25 units)at t = 6: Z = 3.95σ · systematic = 35% of σ²

LOCAL Z AT THIS EXPOSURE3.95 σ

CEILING Z AS EXPOSURE GROWS6.7 σ

EXPOSURE TO REACH 5σ14.3 units

SYSTEMATIC SHARE OF σ²35 %

Live interpretationLOCAL Z AT THIS EXPOSURE: 3.95 σ. CEILING Z AS EXPOSURE GROWS: 6.7 σ. EXPOSURE TO REACH 5σ: 14.3 units. SYSTEMATIC SHARE OF σ²: 35 %

03

Catch the common trap

Explain before calculating.

A bump hunt scans a mass spectrum divided into 250 statistically independent resolution-width windows. The largest excess found anywhere in the scan has a local p-value of 4.0 × 10⁻⁴, which corresponds to 3.35σ locally. What global significance should be quoted?

Choose an answer to test the model.

04

Practice & worked examples

Reason from the model, then test the result.

EasyA counting search expects b = 900 background events in its signal window, with the background known well enough that its uncertainty can be ignored. (a) How many signal events s are needed for a local significance of 5σ? (b) The experiment observes an excess of 60 events. What should it report?
  1. For s ≪ b and large b, the significance of an excess is Z ≈ s/√b, with s and b both event counts.
  2. √b = √900 = 30 events, so a 5σ excess needs s = 5 × 30 = 150 events — a 150/900 = 17% enhancement over the expected background.
  3. The observed excess of 60 events gives Z = 60/30 = 2.0σ. The one-sided probability of a background fluctuation at least this large in one window is 0.023, about 1 in 44.
  4. Report it as a 2.0σ excess, quote the number, and keep running. Two sigma is not evidence; the field reserves 'evidence' for 3σ and 'observation' for 5σ, and 2σ excesses turn up in roughly one search in forty by construction.

Answers = 150 events for a local 5σ; the observed 60 events is a 2.0σ excess (one-sided p ≈ 0.023), to be reported as a fluctuation, not as evidence.

MediumTwo determinations of the Hubble constant: a fit to the cosmic microwave background gives H₀ = 67.4 ± 0.5 km s⁻¹ Mpc⁻¹, and a Cepheid-calibrated supernova distance ladder gives 73.0 ± 1.0 in the same units. (a) Quantify the tension. (b) If the whole discrepancy is an unrecognised systematic in the ladder alone, how large must that systematic be to bring the tension below 3σ?
  1. The two are independent measurements of one quantity, so the uncertainty on their difference adds in quadrature: σΔ = √(0.5² + 1.0²) = √1.25 = 1.118 km s⁻¹ Mpc⁻¹.
  2. Δ = 73.0 − 67.4 = 5.6, so T = 5.6 / 1.118 = 5.0σ.
  3. Five sigma between two measurements of the same quantity is not a discovery of anything. It says that the two error budgets and the model joining them cannot all be right.
  4. To reach T = 3 the combined uncertainty must be at least 5.6/3 = 1.867. Holding the CMB error at 0.5, the ladder's total error must be √(1.867² − 0.5²) = √3.234 = 1.798.
  5. That total is the quoted 1.0 plus a missing term in quadrature: √(1.798² − 1.0²) = √2.234 = 1.49 — an unmodelled systematic of about 1.5 km s⁻¹ Mpc⁻¹, roughly 2% of H₀.

AnswerT = 5.0σ. An unrecognised systematic of about 1.5 km s⁻¹ Mpc⁻¹ — roughly 2% of H₀ — in the distance ladder would bring the tension below 3σ.

HardA direct-detection experiment runs blind: the signal box is opened only after the cuts are frozen. The expected background is 0.20 events for an exposure of 1.0 tonne-year at 60% signal efficiency, and zero events are seen. (a) Give the 90% CL upper limit on the signal rate. (b) A ten-times-larger exposure is proposed, with the background scaling in proportion. Estimate the improvement in the limit.
  1. With no events observed, the Poisson probability of seeing zero is e(−μ), so the classical one-sided upper limit solves e(−μ) = 1 − CL: μᵤₚ = −ln(0.10) = 2.30 events. No background is subtracted at n = 0.
  2. Divide by efficiency and exposure: Rᵤₚ = 2.30 / (0.60 × 1.0 t⋅yr) = 3.8 events per tonne-year at 90% CL. Seeing nothing has measured something — every true rate above 3.8 t⁻¹ yr⁻¹ is excluded.
  3. At 10 t⋅yr with zero background the limit would fall as 1/exposure, to 0.38 t⁻¹ yr⁻¹ — but the background grows with it: b = 0.20 × 10 = 2.0 events.
  4. Observing the expected 2 events, the classical 90% upper limit on the total Poisson mean solves e(−μ)(1 + μ + μ²/2) = 0.10, giving μ = 5.32, so sᵤₚ = 5.32 − 2.0 = 3.3 events.
  5. Rᵤₚ = 3.32 / (0.60 × 10) = 0.55 events per tonne-year — an improvement of 3.83/0.55 = 6.9, not 10. Once b ≫ 1 the limit stops falling as 1/exposure and falls only as 1/√exposure.
  6. Quote it the way the field does: not 'there is no dark matter', but 'rates above this curve are excluded at 90% CL for this mass, for the assumed halo density and velocity distribution'. The assumptions travel with the bound.

AnswerRᵤₚ = 3.8 events per tonne-year at 1 t⋅yr, falling to 0.55 at 10 t⋅yr — an improvement of 6.9, not 10, because the background grew with the exposure.