Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
64 changes: 64 additions & 0 deletions .claude/board/EPIPHANIES.md
Original file line number Diff line number Diff line change
@@ -1,3 +1,67 @@
## 2026-08-12 — E-THE-CONTROL-SCORED-THE-HEADLINE-1

**Status:** FINDING `[G]` — `comet_tail_f16.py` / `.json`, bars committed
BEFORE the run (`05f09005`); report §5.12. EXPLORATORY, not an EV.

**The arc's leading rescue was measured and it failed — and the anti-vacuity
control accidentally calibrated the instrument that had been judging it.**

CT-F16 swapped ONE variable: the dipole's motion reference, from 6h surface
displacement to the 500/600/700 hPa steering flow, on CT-F14's OWN 19 storms
(paired; stored centres reused, so selection and disk geometry cannot move).
Report §9.2 had named this *"the single most promising fix"*. Measured:
**sign consistency 0.579 against a 0.70 bar** (worse than surface's 0.684),
**residual sd 68.29° → 87.71°, 28.4 % WIDER** where a ≥10 % tightening was
predicted, and a level sweep improving **monotonically toward the SURFACE** —
best at 850 hPa, outside the 400–650 hPa band the height ladder predicted,
converging exactly on the surface-displacement figure. There is no
mid-tropospheric optimum on this sample.

**The control is the larger finding.** F16c scored two deliberately WRONG
references through the identical pipeline. The **90°-rotated** steering
reference returned **13/19 = 0.684, p=0.0835** — *numerically identical to
CT-F14's headline*, the number this arc has carried as "suggestive" since
§5.11 (a different set of 13 storms, so the count coincides, not the
identity). At n=19 the one-sided ladder is 11→p=0.324, 13→p=0.0835,
**14→p=0.0318**. So **CT-F14 was never one storm short of significance; it was
one storm short of distinguishability from an answer built to be wrong.**

**Rule: an anti-vacuity control does not only guard the test it is attached to
— it measures the RESOLVING POWER of the instrument.** This one was written to
protect CT-F16 and instead retro-calibrated CT-F14. Attach a
deliberately-wrong reference to any claim whose headline is a rate, and read
the control's score as the floor that headline must clear. Had F16c existed at
§5.11, "0.684, suggestive" would have been reported as "0.684, indistinguishable
from a rotated control".

**Second finding, independent of the first: the sign test conflates a
SYSTEMATIC ROTATION with a correct prediction.** Stratified by steering
strength — weak flow (<10 m/s, n=6) **sign 0.833 / median |err| 103°**; strong
flow (≥10 m/s, n=13) **sign 0.462 / median |err| 55°**; `corr(speed,|err|) =
−0.407`. The prediction gets **more accurate in magnitude** as steering
strengthens, exactly as the physics expects, while sign consistency moves the
**opposite** way. High sign consistency in the weak subset is errors clustered
near −103°: a systematic rotation, which a sign test reports as success. **A
one-sided sign test on a distribution not centred at zero measures which SIDE
the bias falls on, not whether the prediction holds** — and this arc has used
it as the primary instrument since §4. The apparatus story told since §5.9
(*slow storms have noisy bearings → filter on displacement*) is not what these
data show: the well-steered storms are the ones whose signs split at chance
while their magnitudes are best.

**What is NOT falsified.** The height ladder (§5.2/5.8) decomposed the FIELD at
each level about that level's own centre; CT-F16 keeps the SURFACE dipole and
swaps the FLOW reference. Different quantities — the ladder stands as a
measurement (n=2, unreplicated), and what died is the operational reading §9.2
built on it. The structural claim (§9.1: ring profile, wn-1 dominance, the
12-byte carrier) is untouched; nothing in CT-F16 touches it.

**Cross-ref:** `E-ZERO-FOR-ELEVEN-THE-AUTHOR-CANNOT-AUDIT-HIS-OWN-FALSIFIERS-1`
(the author is the wrong person to find a spec's vacuous pass routes — here a
control found a *live* one, in a number already published);
`E-THE-HEADLINE-NUMBER-MEASURED-A-MODEL-NOBODY-CLAIMED-1` (the other time this
arc's headline described something other than what was claimed).

## 2026-08-11 — E-THE-BYTE-WAS-ONLY-THE-SELECTOR-THE-PAIR-IS-THE-CARRIER-1

**Status:** FINDING `[H]` — operator correction + `l4_rail_probe.py` (commit
Expand Down
136 changes: 135 additions & 1 deletion probes/weather-p1/COMET_TAIL_REPORT.md
Original file line number Diff line number Diff line change
Expand Up @@ -770,6 +770,140 @@ this quantity. No stronger attribution is claimed at n = 2.

---

### 5.12 CT-F16 — the steering-level moderator, measured: **it makes the directional claim WORSE**

`comet_tail_f16.py` / `.json`. Bars pre-registered and **committed before the
run** (`05f09005`). §9.2 named this "the single most promising fix" for the
directional claim. It has now been tested and it **fails on both bars.**

**Design.** A *paired* re-scoring of CT-F14's OWN 19 qualifying storms with
**only the motion reference changed** — same storms, same stored centres, same
decomposition, byte-identical `err_deg`. Surface 6h displacement → disk-mean
500/600/700 hPa steering flow. Reusing the stored centres removes storm
selection and disk geometry as variables. **This is a mechanistic test, NOT a
verdict** — a fresh steering-scored sample is CT-F17 and is not run.

| bar | prediction | measured | verdict |
|---|---|---|---|
| **F16a** sign consistency vs steering flow | ≥ 0.70 | **0.579** (11/19), p=0.324 | **FAIL** — and *worse* than surface's 0.684 |
| **F16b** paired residual tightening | sd drops ≥ 10 % | **68.29° → 87.71°, +28.4 % WIDER** | **FAIL** — opposite direction |
| **F16c** anti-vacuity | permuted & rotated both < 0.70 | 0.421 / 0.684 | PASS — but see the caution below |
| **F16d** level sweep (descriptive) | optimum inside 400–650 hPa | **monotone toward the SURFACE; best 850 hPa** | outside the predicted band |

The level sweep is the clearest signal, and it points the wrong way:

| level | 400 | 500 | 600 | 700 | 850 |
|---|---|---|---|---|---|
| sign frac | 0.579 | 0.579 | 0.632 | 0.632 | **0.684** |
| sd (deg) | 89.5 | 87.8 | 82.1 | 80.3 | **77.0** |

Both columns improve **monotonically as the reference level approaches the
surface**, converging at 850 hPa on exactly the surface-displacement figure.
There is no mid-tropospheric optimum. On this sample the surface displacement
was already the best available motion reference, and mid-level flow is a
*worse* one.

**⚠ Does this falsify the height ladder (§5.2/5.8)? No — and the distinction
matters.** The ladder decomposed the **field at each level about that level's
own centre**; CT-F16 keeps the **surface dipole** and swaps only the **flow
reference**. Those are different quantities, so the ladder is untouched as a
measurement. What CT-F16 falsifies is the **application** §9.2 proposed on top
of it — "score the dipole against steering-level motion". That inference is now
measured and dead. The ladder (n=2) remains an unreplicated observation whose
operational reading has failed its first test.

**⚠ The anti-vacuity arm accidentally calibrated the noise floor — and CT-F14's
headline sits exactly on it.** The deliberately **90°-rotated** steering
reference scored **13/19 = 0.684, p=0.0835** — numerically identical to
CT-F14's own headline figure (on a *different* set of 13 storms, so the count
coincides rather than the identity). **A reference constructed to be wrong
produces this arc's "suggestive" number on this sample.** F16c's bar was
`< 0.70` and 0.684 clears it, so the gate passes as written — but the margin is
the finding. At n=19 the one-sided ladder is 11/19 → p=0.324, 13/19 → p=0.0835,
**14/19 → p=0.0318**: the test only separates from chance at 14. CT-F14's
0.684 was never one storm short of significance; it was one storm short of
*distinguishability from a deliberately wrong answer*.

**⚠ And the sign test is not measuring what the arc assumed.** Stratifying by
steering-flow strength inverts the two statistics against each other:

| subset | n | sign frac | median \|error\| |
|---|---:|---:|---:|
| weak flow (< 10 m/s) | 6 | **0.833** | **103°** |
| strong flow (≥ 10 m/s) | 13 | **0.462** | **55°** |

`corr(steering speed, |error|) = −0.407` — the prediction gets **more accurate
in magnitude** as the flow strengthens, which is physically sensible. But sign
consistency moves the *opposite* way. High sign consistency in the weak-flow
subset is not the prediction working: those errors cluster near **−103°**, a
systematic rotation, and a sign test on a distribution centred far from zero
reports the *side* of the offset, not its correctness. The apparatus story this
arc has told since §5.9 — *slow storms have noisy bearings, so filter on
displacement* — is not what these data show; the well-steered storms are the
ones whose signs split at chance while their magnitudes are best.

**Consequence for the arc.** The directional claim does not merely remain
unproven; its **leading mechanistic rescue is now measured and failed**, and
the instrument used to judge it is shown to conflate a systematic rotation with
a correct prediction. §9.2's ranking of the dry moderators is superseded to
that extent: steering level is no longer "the single most promising fix". The
**structural** claim (§9.1) is untouched — nothing here involves the ring
profile, the wn-1 dominance, or the 12-byte carrier.

### 5.13 The instrument was the collapse: circular resultant vs sign test, same 19 storms

`comet_tail_resultant_instrument.py` / `.json`. **Post-hoc re-analysis of
stored rows — explicitly NOT a verdict.** Operator framing: *"die irrationale
Aufsummierung hilft, dass der Dipol nicht auf 0.68 kollabiert."* Measured:

| referent | sign < 0 | R̄ | μ | μ 95 % CI | Rayleigh p |
|---|---:|---:|---:|---:|---:|
| **surface (CT-F14)** | 0.684 | **0.516** | **−30.2°** | ±36.5° | **0.0050** |
| steering (CT-F16) | 0.579 | 0.343 | −40.5° | ±64.3° | 0.107 |
| CONTROL rot+90° | **0.684** | 0.343 | **−130.5°** | ±64.7° | 0.107 |
| CONTROL permuted | 0.421 | 0.142 | — | ±152° | 0.689 |

(uniform-expectation R̄ at n=19 ≈ 0.203)

Three things, in order of importance:

1. **The 0.684 plateau was a property of the STATISTIC, not the data.** The
sign test collapses each storm's error vector to one bit; 19 bits saturate
below the 14/19 distinguishability floor (§5.12). The circular resultant —
the *Aufsummierung*: sum the unit error vectors, read length and direction
— resolves the identical rows at **p = 0.0050**, because concentration (R̄)
and offset (μ) come out as two numbers instead of eating each other. The
systematic ≈−30° offset the arc has chased since §5.1 is now *estimated*
(−30.2° ± 36.5°) instead of *penalizing the score*.
2. **The wrong referent is now visible.** F16c's rotated control scored
0.684 = indistinguishable under the sign test. Under the resultant it has
the same R̄ (rotation preserves concentration, by construction) but μ
shifted **100.3°** — separated by well over both CIs. The instrument
hierarchy is clean: real referent (0.516) > structured-but-wrong (0.343,
wrong μ) > permuted (0.142, below the uniform floor).
3. **NOT a promotion.** Same sample, post-hoc — the p=0.0050 demonstrates the
instrument, it does not establish the directional claim. CT-W6 is the
pre-registered use of circular statistics on these rows (with the
two-component Faltung decomposition); a fresh-sample verdict is CT-F17.

**The Faltung reading (operator, same exchange):** the resultant IS the first
circular Fourier coefficient — a Faltung of the empirical error distribution
with `e^{iθ}`. The generalization is the full harmonic/kernel readout (von
Mises smoothing = circular Faltung; n=19 supports the first 2–3 harmonics),
and the W6 decomposition is a DEconvolution: the measured dipole distribution
= (referent component mix) ⊛ (apparatus noise, ±3–7° measured in CT-F4).
Components add linearly in the transform domain, which is exactly what makes
the multi-referent separation solvable — and on the palette ring `Z_256` the
circular convolution is native to the substrate (`DistanceLut::circular()`'s
own domain, FFT-able at 256 points).

**Consequence for every prior sign-consistency number in this document:**
§4's 2/2, §5.9's 6/10, §5.10's 8/10, §5.11's 13/19 were all read through an
instrument that (a) cannot estimate the offset it penalizes and (b) cannot
distinguish a rotated referent at these n. They stand as recorded, but their
evidential weight is bounded by this section, in both directions — the sign
test neither established the claim nor could it have.

## 6. Product / encoding consequence `[S]`

> **⚠ Read with §5.9–5.11 AND the compression correction in §1.** The figures
Expand Down Expand Up @@ -1119,7 +1253,7 @@ already more explicit structure than a learned model exposes.

| moderator | measured evidence | wiring |
|---|---|---|
| **Steering level** (baroclinic tilt) | the 92–102° monotone height ladder, zero-crossing 400–650 hPa (§5.2/5.8) | score the dipole against the *steering-level* motion (500–700 hPa flow) instead of the 6h surface displacement — the single most promising fix, **CT-F16** |
| ~~**Steering level** (baroclinic tilt)~~ **— TESTED, FAILED (§5.12)** | the 92–102° monotone height ladder, zero-crossing 400–650 hPa (§5.2/5.8), n=2 | ~~score the dipole against steering-level motion — the single most promising fix, **CT-F16**~~ **RUN: 0.579 vs a 0.70 bar, residual 28.4 % WIDER, level sweep monotone toward the SURFACE (best 850 hPa, outside the predicted band). The ladder as a measurement stands; this operational reading of it is dead.** |
| **Displacement magnitude** (label noise) | 6/7 pooled at ≥250 km vs 14/20 unfiltered; CT-F14 0.684 | model the motion-bearing *uncertainty* explicitly instead of a hard cutoff |
| **Surface type / friction** | +14° ocean vs +34° land inflow, paired within one disk (§5.7) | a wind-level correction; second-order on the pressure dipole |
| **Latitude / f, regime** | the low-wn1 July cases; the 75°N outlier | intake covariates, already computed per storm |
Expand Down
Loading