How it works, and how we know

The anomaly log

How a claim here earns its confidence. Every unexpected result — discoveries, certified nulls, refuted repairs, model rejections, and the retractions in between — with the seed that regenerates it. Nothing here is smoothed over: the failed hypotheses are listed with the same typography as the wins, because a record you can only read one way is not a record.

F-0001calibration

ℛ exceeds 1 on matching pennies

ℛ is a norm ratio, unbounded above: values > 1 mean the circulating response dominates the reciprocal part. Docs corrected; the meter was right.

F-0002correction

Only ℛ's zero test is λ-free

The magnitude scales with λ (red-team probe). The λ-free property is the symmetry statement: ℛ = 0 ⟺ potential, at every λ.

F-0004discovery

The meters decouple at high α

Marginal ρ(EPR, ℛ) = 0.993 is α-confounding; within-level coupling collapses in the near-harmonic regime. Found by red-team stratification of our own headline number.

F-0005calibration

Blotto is mixed, not pure harmonic

α = 0.69 on the budget-3 game: zero-sum ≠ harmonic-pure. The realistic high-α anchor sits at ~0.7, not 1.0.

F-0006discovery

Criticality escape + the supercritical wedge

Potential games escape criticality at high λ (concentration kills the choice covariance — now proven a pure λ×scale fold, identity error 0.0); between the anchors a supercritical wedge opens.

▶ play with it — phase map

F-0007refuted repair

Our repair hypothesis, refuted by its own test

H1 predicted the numerator of ℛ would keep tracking dissipation at high α. It doesn't. Response and dissipation are structurally distinct observables — the honest negative that made the instrument-scope table mandatory.

F-0008certified null + retraction

Value-space blindness, and a retracted claim

Price-value discretization is provably blind to loop irreversibility; an intermediate 'certified null' claim was retracted on adversarial review when the null class failed to bracket the data. The retraction is part of the record.

▶ play with it — markets

F-0009detection

The day-ahead market is a driven cycle

Against a persistence-matched reversible null: pair-level detailed balance violated at ~1.1 nats/day (p < 0.01), concentrated in scarcity weeks; verified for null validity, FPR, seeds, bins, ties, multiple testing, and order-2 leakage.

▶ play with it — markets

F-0010failed criterion → finding

Universal collapse, λ-dependent sign

The initial criterion (reversal in ≥3/4 conditions) failed 2/4 — and the failure is the finding: what's universal at α = 0.95 is decorrelation; anti-alignment is a λ-amplified second effect within ~2 null-SD per point.

F-0011empirical read

The reciprocity meter's first real-data read: ℛ ≈ 0.001

Cross-brand wholesale-cost pass-through (Campbell ↔ Progresso, Dominick's scanner panel, 86 stores): the prediction stated in config before the run — one retailer pricing both brands must respond symmetrically — confirmed. Own pass-through 1.07/0.97, asymmetry CI covering zero. A χ row-ordering bug caught pre-review is on the record.

F-0012discovery · mechanism open

The driving cost inverts: pay per change vs pay rent

Quenching λ across the α family: potential games pay only for change (excess ∝ 1/steps, quasi-static driving is free) and nothing to exist; near-harmonic games pay almost nothing for change but burn constant housekeeping rent just to hold their steady state. The first-pass mechanism was refuted by the red-team's own probe — the collapse is real, its cause is open.

F-0016withheld → certified

The estimator that had to earn it

The data-facing quench meter was refused by adversarial review: its natural self-check (the fluctuation theorem) proved structurally insufficient — a 45% bias hid behind an IFT of 1.01. The chase refuted two hypotheses, found a real missing-window bug, vindicated the statistics (20/20 calibration), replaced the check with a physical relaxation gate — and a fresh review granted certification with every escalation on the record.

F-0019refusal with a mechanism

The number survives a month of data; the verdict doesn't

Can a month of market data (~30 trajectories) be quoted against a certified floor of 200? No — but not for the expected reason. Interval coverage holds all the way down to n = 20; what fails is the instrument's own decision machinery, which needs roughly ten times more data than the estimate does. A permutation of the trajectories — which cannot change any physical property — was flipping the verdict, and that localised the blame to one implementation choice rather than to the physics.

F-0020fixed → partly retracted

Two failure modes that looked like one

The suspect was an arbitrary 4-way split of the trajectories, used to estimate an error bar. Replacing it with order-invariant estimators kills the instability completely — zero flips at every n, and on real price data the flip rate falls 0.214 → 0.050 with one month collapsing 17/20 → 0/20. It even flips a real month from refused to admitted while leaving the estimate identical to six decimals. But the small-n floor does not move, which retracts half of what the previous finding predicted: the relaxation-time estimate itself varies 35–40% across seeds at n = 30, so an error bar that is accurate must report that variance — and must therefore refuse holds that are genuinely settled. Stability was a bug; accuracy is a limit. The same permutation trick then exposed a second, independent order-dependence hiding in the interval bootstrap, where the obstruction is structural: a yes/no flag thresholded on a Monte-Carlo interval sitting at its threshold cannot be stabilised by any choice of seed.

▶ try it · F-0004/F-0010 · you predict it
in-browser · goldens-checked

Within each α level, how does the correlation between dissipation (EPR) and response asymmetry (ℛ) behave as games get more harmonic? Commit to a guess — rises, flat, or collapses — then reveal the measured curve.

▶ try it · F-0006 · the supercritical frontier λ_c(α)
in-browser · goldens-checked
λ_c — first crossing of ρ = 1λ_c (median game)0.30.8

The wedge boundary, measured by verified-single-crossing bisection: more harmonic content ⇒ criticality arrives at lower λ. Below α ≈ 0.5 the median game never crosses by λ = 15.

▶ try it · F-0011 · the empirical pass-through matrix
committed artifact · seeded

Poke Campbell's wholesale cost, read both shelf prices; poke Progresso's, read both again. That is the poke panel's procedure run on 22,655 real store-weeks. One retailer prices both brands — so before revealing: should the two cross-readings agree (a landscape) or disagree (a whirlpool)?

∂ log p / ∂ log cCampbell costProgresso cost
Campbell price??
Progresso price??
▶ try it · F-0012 · pay per change: excess dissipation vs α
committed artifact · seeded
log₁₀ nats per ramp (λ 0.5 → 3.0)log₁₀ excess ⟨Y⟩00.95

The cost of *changing* λ collapses three orders of magnitude toward the harmonic end — and honesty note: the first-pass explanation for why was refuted by the red-team's own probe. The collapse is measured; its mechanism is an open chase item, on the record.

▶ try it · F-0012 · pay rent: housekeeping vs α
committed artifact · seeded
housekeeping over the same ramp∫σ_hk dt (nats)00.95

The fuel burned just to *hold* the steady state: exactly zero for potential games, 12.4 nats at α = 0.95 — at a constant 0.774 nats per unit time, charged whether you drive or not. Quasi-static driving is free only on a landscape.

▶ try it · thermo.estimators · data-side meters vs the exact one
in-browser · goldens-checked
nats / unit time across the α familyexact EPRKLD estimateTUR certified0.050.95

The KLD estimator sits on the exact meter (ρ = 1.0); the TUR certified bound stays below it everywhere, tightest near equilibrium — the calibration that made the market reading of F-0009 trustworthy.