From ffef98f95dfd8697bdf621213beb62418420a9d0 Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 25 Aug 2026 16:28:32 +0000 Subject: [PATCH 1/6] board+plan: the transfer finding, and R2IL-as-machine-semantic-contract v1 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two artifacts, one commit, per the board-hygiene rule. EPIPHANIES: E-R2IL-MACRO-VOCABULARY-TRANSFERS-ACROSS-COMPILER-AND-LANGUAGE-1 — the 33-macro vocabulary trained on two gcc binaries fires in an unseen gcc program at -0.6% density and in an unseen rustc binary at -4.7%, both outside a column-marginal-preserving shuffle null (20 seeds, multiset-asserted per draw). Deflations stated in the entry: 33/33 is partly sample size, the Rust harvest is capped at 200/548 functions, coverage saturates and density carries the finding. The entry also records that my own pre-registration missed in the hypothesis's favour. Plan: r2il-machine-semantic-contract-v1.md — answers the sibling session's 'how should R2IL be stored' question in five rules derived from shipped code (no private object graph; a V4 tenant physically identical to V3; the 12-byte register carved per le-contract section 3 and read by reinterpret; a shared palette, not FnIndex-per-macro; behaviour by address). Separates documentation (cited) / plan (falsified) / status (D-ids) as registers, reproduces the operator's dependency-layering sketch with its correction (r2sleigh uses R2IL as its contract; R2IL is backed by SoA V4; libsla stays as the oracle and is removable LAST), and sequences six waves each with a kill condition, W0 being the Custom(n) content-vs-reading census that gates the 0xC4 mint. --- .claude/board/EPIPHANIES.md | 103 +++++++++ .../r2il-machine-semantic-contract-v1.md | 211 ++++++++++++++++++ 2 files changed, 314 insertions(+) create mode 100644 .claude/plans/r2il-machine-semantic-contract-v1.md diff --git a/.claude/board/EPIPHANIES.md b/.claude/board/EPIPHANIES.md index 7acb7b92c..06d22ab76 100644 --- a/.claude/board/EPIPHANIES.md +++ b/.claude/board/EPIPHANIES.md @@ -1,3 +1,106 @@ +## 2026-08-25 — E-R2IL-MACRO-VOCABULARY-TRANSFERS-ACROSS-COMPILER-AND-LANGUAGE-1 — a macro vocabulary learned from two gcc binaries fires in unseen gcc code at −0.6% density and in unseen rustc code at −4.7%, both outside a marginal-preserving null + +**Status:** FINDING — [MEASURED] (held-out transfer + two pre-registered +shuffle nulls, 20 seeds each, over 4 binaries in 3 corpus configurations; +the shipped `probe_bpe_r2il_loco_microcode.rs` instrumented for the runs +and reverted — no source change landed, the probe's own 10 gates stayed +green throughout). +**Confidence:** High for the mechanism on this corpus. The generalization +beyond x86-64 / the pass-1 seven-opcode convention / chain-length-3 is +NOT measured and is not claimed. + +### The question this answers + +`E-BPE-OVER-DEFUSE-CHAINS-BEATS-LINEAR-AND-FITS-LOCO-1` measured 33 BPE +macros over R2IL def-use chains and their `ogar-loco` lane fit, on **2 +binaries**. It did not ask whether a macro means the same thing anywhere +else. That question is load-bearing for any design that treats a macro as +a placeable, addressable unit: a unit whose meaning is program-local is +content, not vocabulary. + +### The measurement + +Train on the two `stress_test` binaries ONLY (1,872 def-use chain +occurrences, 33 macros). Score each unseen corpus by macro-hit **density +per chain** — scale-free, so corpora of different size compare directly. + +| held-out | chains | density | vs train | coverage | macros hit | +|---|---|---|---|---|---| +| *(train: 2 × gcc `stress_test`)* | 1,872 | **2.529** | — | — | — | +| `vuln_test` — different C program, same gcc | 608 | **2.515** | **−0.6%** | 608/608 | 31/33 | +| `build-script-build` — serde_json, rustc/LLVM | 8,659 | **2.409** | **−4.7%** | 8,652/8,659 | 33/33 | + +### Both nulls, pre-registered before the runs + +- **Global null** — permute atoms ACROSS held-out chains; the global atom + multiset is preserved EXACTLY (asserted on every one of the 40 draws), + def-use adjacency destroyed. +- **Column null** (the one that does the work) — permute each POSITION + column among itself. Each position's opcode marginal is preserved + exactly, so *"position 0 is usually `int_add`"* cannot explain a win; + only which x, y and z sit TOGETHER is destroyed. + +| held-out | REAL | global null | column null | REAL in either range? | +|---|---|---|---|---| +| `vuln_test` | 2.515 | 1.664 [980..1034 hits] | 1.969 [1179..1232 hits] | **no / no** | +| Rust | 2.409 | 1.780 [15278..15496] | 2.009 [17327..17461] | **no / no** | + +Margin over the strict null: **+27.7%** (same compiler) → **+19.9%** +(across the language boundary). The cost of crossing is real and small; +the vocabulary is NOT a gcc idiom. + +Distribution-free ranges are reported deliberately: `I-NOISE-FLOOR-JIRAK` +forbids a classical σ claim here (weakly dependent bits), and +non-overlap of 20-draw ranges needs no such assumption. + +### ⊘ My own pre-registration was WRONG, in the hypothesis's favour + +Before the Rust run I recorded: *"Erwartung: schwächerer Transfer … +spürbarer Abfall der Dichte"*, and named the kill condition (REAL inside +the column-null range ⇒ the palette must be split per toolchain). +Measured: 4.7%, an order of magnitude smaller than "spürbar" implied, and +the kill condition did not fire. Recorded because a prediction that misses +in the direction you wanted is exactly the one a later reader must be able +to check. + +### Three deflations of the headline, stated rather than buried + +1. **33/33 is partly a sample-size effect, NOT a stronger transfer than + 31/33.** The two macros that missed `vuln_test` carry 6 and 2 training + occurrences and had 608 chances there against 8,659 here. Density is + the honest statistic, and it correctly shows the Rust corpus as the + HARDER one. +2. **The Rust sample is capped.** `R2IL_HARVEST_MAX_FUNCS=200` exhausted + at 200 of 548 `STT_FUNC` symbols — a first-200 slice with std/core + prelude bias. The three C binaries sat under the cap and were harvested + whole, so the four probes are not sampled identically. +3. **Coverage saturates and should not be quoted alone.** 608/608 looked + tautological for a 7-symbol alphabet; the column null reaches only + 564/608 (92.8%), which is what rescues it — but the margin is thin and + density carries the finding. + +### What this licenses, and what it does not + +**Licenses:** treating the R2IL macro vocabulary as a *shared* palette +rather than a per-binary or per-toolchain one — the `System[256] | +Learned[≤256] | Explore[≤256]` shape of the POC entry, with one mint +serving multiple programs. It is the empirical half of "R2IL is the +faithful vocabulary, BPE is recombination over it". + +**Does not license:** any mint (none performed), any claim about +architectures other than x86-64, any claim that a macro is *semantically* +the same across languages — this measures co-occurrence of an opcode +pattern, not that the pattern means the same thing to a reader. A +`(int_add, copy, store)` chain in Rust and in C are the same SHAPE; that +they are the same THOUGHT is a separate, unrun question. + +**Corpus:** `stress_test` (10,003 rows) · `stress_test_opt` (7,554) · +`vuln_test` (4,409, 41 fns) · `build-script-build` (72,567, 200 of 548 +fns). Harvested through `ruff_r2il`'s `harvest_r2il` (`--features lift`). +The probe's B7 fence — which recomputes the binary count live rather than +trusting a constant — correctly failed the moment a third binary entered, +which is how the corpus swap was confirmed real. + ## 2026-08-24 — E-R2IL-BPE-RECOMBINATION-FALSIFIERS-CONFIRMED-1 — the typed genetic recombination proposal's three §7 falsifiers all run green: splice points exist selectively (10.1%), recombination round-trips, and the counterfactual lane distinguishes at the real (v2-only) primitive level **Status:** FINDING — [MEASURED] (`PROBE-R2IL-BPE-RECOMBINATION-FALSIFIERS-1`, diff --git a/.claude/plans/r2il-machine-semantic-contract-v1.md b/.claude/plans/r2il-machine-semantic-contract-v1.md new file mode 100644 index 000000000..7d43dd4e3 --- /dev/null +++ b/.claude/plans/r2il-machine-semantic-contract-v1.md @@ -0,0 +1,211 @@ +# R2IL as the machine-semantic contract — documentation, plan, status, integration + +> **Status:** PLAN (v1, 2026-08-25). One measured finding underneath +> (`E-R2IL-MACRO-VOCABULARY-TRANSFERS-ACROSS-COMPILER-AND-LANGUAGE-1`); +> everything else here is PROPOSED and labelled as such, line by line. +> **Nothing in this document has been built.** No mint performed, no +> layout bump, no crate created. +> +> **Origin:** an operator sketch of the dependency layering (§2), plus a +> sibling session's live question — *"wie soll die Session R2IL +> speichern?"* — answered in §4. This plan exists so that question is +> answered once, from shipped code and measurement, instead of +> re-derived per session. + +--- + +## §0 — The three registers, and why they are separate + +The single most common failure in this workspace is a sentence that is +half description and half proposal, so a later reader cannot tell which. +This plan keeps them apart mechanically, and the separation is the +deliverable as much as the content is: + +| register | answers | where it lives | rule | +|---|---|---|---| +| **DOCUMENTATION** | *what IS, today, in shipped code* | §3 here; module docs; `docs/` | every claim carries `file:line` or a measured number. If it cannot, it is not documentation | +| **PLAN** | *what is PROPOSED, and what would falsify it* | §5–§6 here; `.claude/plans/` | every item carries a falsifier and a kill condition. A plan item with no way to fail is a wish | +| **STATUS** | *where we actually are* | `.claude/board/STATUS_BOARD.md` | D-id + one of Queued / In progress / In PR / Shipped / Closed. Never prose | +| *(evidence)* | *what was measured, once* | `.claude/board/EPIPHANIES.md` | append-only; a finding is cited from here, never restated | + +**The test, applied to this very document:** §2 is an operator sketch — +it is neither documentation nor measurement, and it is marked as such. +§3 is documentation and every row cites source. §4 is an ANSWER derived +from §3 + the evidence, so it inherits their status, not a higher one. +§5 onward is plan. + +--- + +## §2 — The dependency layering (OPERATOR SKETCH, not yet built) + +Stated by the operator 2026-08-25. Reproduced because its *correction of +an earlier framing* is the load-bearing part: + +``` + GHIDRA + UI / plugins / analyzers / API (Java, unchanged) + │ consumes a Java compatibility surface + ▼ + lance-graph-java (Valhalla descriptors / views) + │ Panama FFM — crossed ONCE + ═════════════════════╪═════════════════════════════════ + ▼ + r2sleigh (SLEIGH/P-code → R2IL translator) + │ consumes the r2il crate + ▼ + R2IL (faithful machine-semantic ISA) + │ lowers onto + ▼ + SoA V4 (lanes / masks / SIMD / ClassView) + │ + lance-graph +``` + +**The correction, in the operator's own terms.** An earlier framing had +`r2sleigh` owning the native representation: + +``` + WRONG: Ghidra Java → Panama → r2sleigh → "R2IL SoA" + ^^^^^^^^ owns a private object graph, + serializes some R2IL later + + RIGHT: r2sleigh USES R2IL as its machine-semantic contract, + and R2IL itself is backed by SoA V4. +``` + +Consequence: **r2sleigh does not invent a second storage physics.** It +becomes primarily the SLEIGH/P-code → R2IL translator. That is not a +stretch — it is what its own README already describes (§3). + +**The Ghidra half — compatibility FACADE, not compatibility DATA MODEL.** +`instruction.getPcode()` keeps working, but the implementation becomes +`InstructionRef → block handle → SoA slice/mask → lazy materialization`. +The cost of manufacturing `PcodeOp[]` is paid ONLY when old Java asks for +it; new analysis operates on masks and blocks and never pays. This is the +same doctrine `lance-graph-java`'s own mask-native policy already states +for `where`/`hop`/`compute` — Java-surface convenience never dictates +substrate representation. + +--- + +## §3 — DOCUMENTATION: what ships today (every row cited) + +| claim | source | +|---|---| +| r2sleigh's own pipeline is `.sla (Ghidra) → libsla → P-code → r2il → {ESIL, SSA, decompiler, type inference}` | `r2sleigh/README.md` | +| `r2il` is the typed IL (60+ opcodes); `r2sleigh-lift` is the SLEIGH/P-code translation layer | `r2sleigh/README.md`, `crates/{r2il,r2sleigh-lift}` | +| **`libsla` is a dependency of exactly ONE crate** | `crates/r2sleigh-lift/Cargo.toml:10` — the whole workspace's Ghidra-native coupling is isolated to that line | +| `ruff_r2il` consumes `r2il`/`r2ssa` by path and melts them into flat facet-addressed rows | `ruff/crates/ruff_r2il/{Cargo.toml:19, src/lib.rs}` | +| The R2IL address today is `VarnodeFacet` = 16 bytes, `classid \| offset_lo \| offset_hi \| size` | `ruff_r2il/src/facet.rs:40,232` | +| The 12-byte content-blind register carving is `CascadeShape::{G6D2,G4D3,G3D4}`, all `G·D=12` | `lance-graph-contract/src/facet.rs:359-431` | +| …and is deliberately MIRRORED in OGAR as `LaneShape::{Pairs,Triples,Quads}` | `OGAR/crates/ogar-loco/src/lib.rs:1-100` ("mirroring the LE contract's CascadeShape"); graded ESTABLISHED by the LG #1023 audit | +| `FacetCascade` decode is a pointer reinterpret, not a copy | `facet.rs:93` `#[repr(C, align(16))]`; `as_bytes`/`ref_from_bytes`; test `reinterpret_is_a_no_op` asserts POINTER IDENTITY | +| …and that zero-copy path has **zero non-test consumers** — every real caller uses the copying `from_bytes` | measured 2026-08-25: `ref_from_bytes` appears nowhere outside `facet.rs` | +| The macro vocabulary space is `System[256] \| Learned[≤256] \| Explore[≤256]`, palette-indexed per lane | `lance-graph-contract/src/soa_view.rs:41`; `E-BPE-OVER-DEFUSE-CHAINS-BEATS-LINEAR-AND-FITS-LOCO-1` | +| R2IL's real opcode census is 9 RISC-like opcodes — **not** a bitmask ALU | LG #1023 audit Challenge 1, 143 fns / 17,557 rows | +| "IR lands in V4, physically identical to V3 — no `ENVELOPE_LAYOUT_VERSION` bump" | operator-stated, recorded in the POC entry | +| `CONTENT NEVER TRAVELS IN CLASSID. CLASSID SELECTS THE READING.` | ratified PR #1012, `E-CONTENT-NEVER-TRAVELS-IN-CLASSID-1` | + +**One documented defect, found 2026-08-25 and not yet filed as work.** +`facet.rs:232` computes `classid = (concept << 16) | space_discriminant`. +For the FIXED spaces (`Register`/`Ram`/`Const`/`Unique`) that is +legitimate — the space genuinely selects how the remaining 12 bytes are +read, which is exactly what a classid is for. For **`SpaceId::Custom(n)`** +it is not: `n` is a per-binary ordinal lifted out of the program, i.e. +content in the reading selector, against both the ratified law above and +OGAR's canon (*"lo u16 = APP render prefix — NEVER a shape ordinal"*). +The open item O3 currently reads *"does `Custom(u32)` fit the 16-byte +projection?"*; under this reading the question is not whether it fits but +that it does not belong. **CONJECTURE** — the falsifier is in §6. + +--- + +## §4 — THE ANSWER: how a session should store R2IL + +The sibling session's question, answered from §3 + the evidence, in five +rules. Status inherits from the sources: rules 1–3 are documentation, +rules 4–5 are plan. + +**R1 — Do not build a private object graph and serialize R2IL later.** +That is the exact shape the operator's correction rejects, and the +Firewall (ADR-022/023) forbids serialization in the hot path regardless. +`to_le_bytes`/`from_le_bytes` ARE the format. + +**R2 — R2IL lands as a V4 tenant that is PHYSICALLY V3.** 16-byte +content-blind facet, `classid` selects the reading, no +`ENVELOPE_LAYOUT_VERSION` bump. V4 is an additive sibling tier for R2IL's +100%-coverage requirement (operator ruling 2026-08-21, +`E-V4-IS-THE-100-PERCENT-TIER-V3-UNCHANGED-1`); V3 continues unchanged +for everything it already carries. + +**R3 — The 12 bytes are a dumb register the ClassView projects.** Carve +per `le-contract.md` §3 — `6×(8:8)` rails / `4×(8:8:8)` / `3×(8:8:8:8)`. +`u8:u8` stays two separate bytes, never widened. Reads go through +`ref_from_bytes` (a reinterpret), not `from_bytes` (a copy) — see the +§3 row saying nobody currently does this. + +**R4 — The macro vocabulary is a PALETTE, not a `FnIndex` per macro.** +`ogar-loco` ROUTES microcode into the `System`/`Learned`/`Explore` +palettes; the "90 free slots" number is one encoding's headroom, not a +vocabulary ceiling. §5's evidence says one palette can serve several +programs and two toolchains, so the palette is shared, not per-binary. + +**R5 — Behaviour travels by ADDRESS, never inline.** The macro id is an +ordinal into a class's set; the class's `ClassView` resolves it. An +inline lambda on the surface is the SURREAL-AST trap in its R2IL edition. + +--- + +## §5 — THE EVIDENCE that makes the shared palette defensible (2026-08-25) + +`E-R2IL-MACRO-VOCABULARY-TRANSFERS-ACROSS-COMPILER-AND-LANGUAGE-1`, in +one table. Trained on the two gcc `stress_test` binaries only; scored by +macro-hit density per def-use chain (scale-free); two pre-registered +nulls, 20 seeds each, multiset assertions on every draw: + +| held-out corpus | density | vs train (2.529) | strict (column) null | REAL in range? | +|---|---|---|---|---| +| unseen C program, same gcc | 2.515 | −0.6% | 1.969 [max 1232 hits vs real 1529] | no | +| unseen Rust (rustc/LLVM) | 2.409 | −4.7% | 2.009 [max 17461 vs real 20861] | no | + +The vocabulary is not a gcc idiom; crossing the language boundary costs +4.7%, not a collapse. The pre-registered kill condition (REAL inside the +column-null range ⇒ split the palette per toolchain) did not fire. Fences +carried, not waived: x86-64 only, pass-1 seven-opcode convention, +chain-length 3, the Rust sample capped at 200/548 functions. + +**Why this matters to §4/R4:** a palette entry is only mintable as SHARED +vocabulary if it means something outside the binary it was learned from. +That is now measured on the axes available in this workspace. What it +does NOT establish: that the shape is the same THOUGHT across languages — +co-occurrence is measured, semantics is not. + +--- + +## §6 — PLAN: the wave order, each with its falsifier + +No wave starts before its predecessor's falsifier is green. Model policy +per workspace rule (grindwork → Sonnet; accumulation/orchestration/gates +→ Opus; gates run centrally, once). + +| wave | deliverable | falsifier / kill condition | +|---|---|---| +| **W0** | `Custom(n)` census (closes §3's defect + O3): over the 4-binary corpus, count distinct `Custom(n)` and whether the same `n` denotes the same thing across binaries | same `n`, different meaning across binaries ⇒ `Custom(n)` is CONTENT ⇒ it moves OUT of classid into payload/edge before ANY mint. Same meaning everywhere ⇒ reading, stays | +| **W1** | the R2IL V4 tenant spec: ClassView carving of the 12 bytes for Op / Varnode-ref / macro-ref rows, written against `le-contract.md` §3 | field-isolation matrix test (I-LEGACY-API-FEATURE-GATED): write each field, assert all others unchanged. Any aliasing ⇒ re-carve | +| **W2** | the `0xC4 BinaryLifting` mint (ruff PR3 arc, O5) — container concepts only, gated on W0's verdict for the space axis | mint request names concepts; a concept that encodes per-binary content is rejected by W0's rule | +| **W3** | palette wiring: macro-id byte → `{Learned,Explore}Palette[id]` → R2IL×BPE microcode, `ogar-loco` routing | B4-equivalent byte-exact round-trip through the palette; admission stays MUL's/the triangle's, NEVER the wiring's | +| **W4** | `r2sleigh` seam: `r2sleigh-lift` emits INTO the tenant (R1 — no private graph then serialize). libsla STAYS — it is the oracle | parity: lift the same binary via today's path and via the tenant; byte-equal FlatFact streams. libsla removal is LAST (see below), never first | +| **W5** | `lance-graph-java` facade: `PcodeOp`/`Instruction`/`Varnode` as lazy views over handles + masks | the mask-native gates (GraphHopTest allowlist, allocation gates) extended to the new views; `getPcode()` pays materialization ONLY when called — measured, not asserted | + +**libsla is removable scaffolding, and deliberately LAST.** While it is +in place, both sides read the same SLEIGH truth, which makes it the +byte-parity oracle for every wave above (the tesseract-rs method). It is +removed only after W4's parity is green — the "final removable +scaffolding", never an early surgery. This also keeps ruff PR4 +("libsla-optional") correctly sequenced: it is a *candidate* after W4, +not a prerequisite of anything. + +**Open questions this plan does NOT decide** (inherited from the V4 +ruling, still with the operator): V4's exact special-needs set beyond +R2IL; the V3-vs-V4 routing rule; whether the W2 mint is registered as V3 +or V4. From c43071f8f6f9d9df3198df41d8676747fe0be925 Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 25 Aug 2026 16:34:21 +0000 Subject: [PATCH 2/6] =?UTF-8?q?plan(r2il-contract):=20section=207=20?= =?UTF-8?q?=E2=80=94=20white/grey=20vocabulary,=20hex=20demoted=20to=20a?= =?UTF-8?q?=20testable=20overlay,=20demotion=20gate?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Records the converged framing as plan-register material, with the discipline the document itself demands: the six-layer vocabulary names only shipped structure (each line marked shipped/measured/unbuilt); the hexagonal reading stays FALSE/net-new for storage per the #1023 audit and is re-proposed strictly as an optional learned neighbourhood overlay with a defined A/B probe (metrics + kill condition), plus the honest gate that its cheap first version runs on the existing 33-macro co-occurrence graph before any new tissue is built; and crystallization (Explore->Learned->System) gains its measured admission shape from today's transfer finding plus the missing demotion path, with a two-sided can-fire / can-stay-silent falsifier, so the vocabulary cannot become a guard that only ever grows. --- .../r2il-machine-semantic-contract-v1.md | 78 +++++++++++++++++++ 1 file changed, 78 insertions(+) diff --git a/.claude/plans/r2il-machine-semantic-contract-v1.md b/.claude/plans/r2il-machine-semantic-contract-v1.md index 7d43dd4e3..1f0a7be67 100644 --- a/.claude/plans/r2il-machine-semantic-contract-v1.md +++ b/.claude/plans/r2il-machine-semantic-contract-v1.md @@ -209,3 +209,81 @@ not a prerequisite of anything. ruling, still with the operator): V4's exact special-needs set beyond R2IL; the V3-vs-V4 routing rule; whether the W2 mint is registered as V3 or V4. +--- + +## §7 — ADDENDUM 2026-08-25: the white/grey reading, hex demoted to a testable overlay, and the demotion gate + +> Register: PLAN + vocabulary. One operator-converged framing, recorded +> because it names shipped structure correctly and turns the one unbacked +> piece (hexagonal topology) from a rejected reading into a falsifiable +> experiment. Nothing here changes W0–W5. + +### 7.1 The layer vocabulary (naming, not new machinery) + +``` +PHYSICS SoA V4 (shipped) +SEMANTICS R2IL contract (shipped; ruff_r2il consumes it) +ADDRESSING Cartesian / Morton / HHTL (shipped; 256×256 per rail, 4⁴ centroids) +HOLOGRAPHY bounded VSA superposition (shipped, NICHE: N ≤ √d/4 ≈ 32, + I-VSA-IDENTITIES; not a storage story) +VOCABULARY transferable BPE macros (measured 2026-08-25: −0.6% gcc→gcc, + −4.7% gcc→rustc, both outside the + marginal-preserving null) +COGNITION learned topology + resonance (System/Learned/Explore lanes shipped; + the TOPOLOGY half is UNBUILT) +``` + +White matter = PHYSICS…VOCABULARY's exact half (deterministic, addressed, +replayable; `expand(macro) → exact R2IL` is B4, already green). Grey +matter = the plastic half (resonance, counterfactual lane, revision — +`PROBE-METACOGNITIVE-TRIANGLE-1`, PR #1013's veto). The bridge invariant, +already ratified in the POC entry: **a macro never becomes a second +truth.** + +### 7.2 Hexagonal topology — demoted correctly, and thereby testable + +The #1023 audit's FALSE/net-new verdict on "6 directional 16-bit facets" +stands UNTOUCHED: hex is **not** the storage geometry, and no rereading +of the 96-bit register claims otherwise. The corrected proposal is an +**optional learned neighbourhood graph ON TOP of the Cartesian crystal**: + +``` +Cartesian resident crystal → addressed cells/masks/edges + → optional learned neighbourhood overlay → 6-neighbour resonance +``` + +**PROBE-GREY-TOPOLOGY-AB (unbuilt, gated).** Two topologies, same macros, +same tasks, same activation rules: + A = Morton/Cartesian neighbourhood (the substrate's native reading) + B = six-neighbour graph overlay +Metrics: pattern completion, transfer, steps-to-convergence, false +resonance, memory footprint, counterfactual quality. +Kill condition: B must beat A on ≥2 metrics without losing on footprint +by more than it gains — otherwise the crystal stays square. + +**Honest gate:** the experiment as specced requires a learned-topology +layer that does not exist. The cheapest real first version needs no new +tissue: the 33-macro co-occurrence graph over the existing 4-binary +corpus IS a learned neighbourhood — run completion/false-resonance on it +under both adjacency readings before building anything. If even that +cheap form shows no B-advantage, the full probe is not worth its cost. + +### 7.3 Crystallization has a measurable admission criterion now — and needs a DEMOTION gate + +Admission (Explore→Learned→System) stays MUL's + the triangle's. What +2026-08-25 adds is the evidence shape: **a System candidate must +(a) B4-round-trip, (b) fit a loco lane, (c) fire across program AND +toolchain boundaries outside the marginal-preserving null.** On today's +corpus: 25 macros qualify, 7 stay lane-local, 2 (`fits-NONE`, both +binary-local) never crystallize. + +**The missing half, required before any promotion ships:** the demotion +path. A System macro that falls INSIDE the null range on a new corpus +goes back to Learned — same gate, run in reverse. Without it, +crystallization is the 150/150 guard one level up +(`E-ANTI-EIGENVALUE-…-1`): a vocabulary that can only grow carries +progressively less information per entry. Promotion and demotion are one +mechanism read in two directions, and the falsifier for the pair is: +inject a deliberately corpus-local macro into System, present the new +corpus, assert it is demoted (can-fire) — and assert the 25 transferable +ones are NOT (can-stay-silent). From eb4486865c30f8a3dbf54b67296275ab694c1bc5 Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 25 Aug 2026 16:34:52 +0000 Subject: [PATCH 3/6] =?UTF-8?q?plan(r2il-contract):=207.4=20=E2=80=94=20th?= =?UTF-8?q?e=20reversible-trust=20state=20machine,=20operator-converged?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit System is long-term memory, never truth; truth stays with white matter (CLAIM boundary, B4 expansion, contract veto). Crystal is a reversible trust STATUS, with symmetric promote/demote criteria under one epistemic yardstick, required hysteresis (promote threshold above demote threshold, anti-flutter not leniency, values to be measured as policy pins), and the sharpened hex burden of proof: beat Morton measurably as a compute topology or leave — never a retroactive reading of the register. --- .../r2il-machine-semantic-contract-v1.md | 52 +++++++++++++++++++ 1 file changed, 52 insertions(+) diff --git a/.claude/plans/r2il-machine-semantic-contract-v1.md b/.claude/plans/r2il-machine-semantic-contract-v1.md index 1f0a7be67..7ac3d3fd0 100644 --- a/.claude/plans/r2il-machine-semantic-contract-v1.md +++ b/.claude/plans/r2il-machine-semantic-contract-v1.md @@ -287,3 +287,55 @@ mechanism read in two directions, and the falsifier for the pair is: inject a deliberately corpus-local macro into System, present the new corpus, assert it is demoted (can-fire) — and assert the 25 transferable ones are NOT (can-stay-silent). + +### 7.4 Operator refinement, same day: the reversible-trust state machine + +Converged wording, recorded verbatim in substance because it is the +demarcation the rest of the plan leans on: + +> **White Matter ist Wahrheit und Zwang. Grey Matter ist Hypothese und +> Plastizität. Crystal ist der reversible Vertrauensstatus gelernter +> Muster. Hexagon ist noch gar nichts außer einem Kandidaten für lokale +> Rechengeometrie.** + +`System` therefore means *"currently confirmed strongly enough to be +treated as a fixed palette"* — long-term memory, never truth. Truth is +carried by white matter only (the CLAIM boundary, B4 expansion, the +contract veto). Crystal is a STATUS of a pattern, not a third storage +technology. + +``` + promote promote +Explore ───────────────→ Learned ───────────────→ System + ↑ ↑ │ + └──────── demote ────────┴──────── demote ────────┘ +``` + +**Symmetric criteria, one epistemic yardstick in both directions:** + +``` +PROMOTE if B4 exact round-trip + AND fits loco geometry + AND transfer density stable + AND REAL outside the marginal-preserving null + AND the contract falsifier passes + +DEMOTE if B4 breaks + OR the contract falsifier fails + OR a new held-out corpus falls into the null range + OR transfer density collapses beyond admitted tolerance +``` + +**Hysteresis is REQUIRED, not optional:** `promote threshold > demote +threshold` — so a single exotic binary cannot flap the state machine at +a boundary case. This is an anti-flutter measure, not leniency; the +tolerance values are POLICY PINS to be measured when the first promotion +ships, not defended. + +**Hex burden of proof, sharpened.** The question is never "can a hexagon +do association" (the shipped systems already do). It is: *can a +6-neighbour topology beat the existing Morton/Cartesian structure +measurably* — same macros, same corpus, same promotion rules, same +resonance rule. Metrics extend §7.2's list with: promotion stability and +demotion rate. If hex wins, it wins as a COMPUTE topology — never +retroactively as an explanation of the 96-bit register. From 535001ec4e5a5d1eafc18d47ef81aa59d7963a3b Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 25 Aug 2026 16:37:56 +0000 Subject: [PATCH 4/6] =?UTF-8?q?plan(r2il-contract):=207.5=20=E2=80=94=20ca?= =?UTF-8?q?libration-before-demotion,=20operator-hardened?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A corpus may falsify a crystal only after it has falsified the hypothesis that it is an incompetent instrument. Records the defect in my own first draft (the established macros deciding whether the corpus may demote them — a self-immunizing loop) as a correction rather than a silent fix, then the six-point corpus qualification with both control directions (known signal seen, known noise not seen), the leave-one-out form that lets an established macro actually die, the frozen CAL+/CAL- panel target (B4's corruption demo named as CAL- seed material), and demotion-by-accumulated-evidence under the 7.4 hysteresis so one strange binary cannot empty the library. --- .../r2il-machine-semantic-contract-v1.md | 76 +++++++++++++++++++ 1 file changed, 76 insertions(+) diff --git a/.claude/plans/r2il-machine-semantic-contract-v1.md b/.claude/plans/r2il-machine-semantic-contract-v1.md index 7ac3d3fd0..038a3633d 100644 --- a/.claude/plans/r2il-machine-semantic-contract-v1.md +++ b/.claude/plans/r2il-machine-semantic-contract-v1.md @@ -339,3 +339,79 @@ measurably* — same macros, same corpus, same promotion rules, same resonance rule. Metrics extend §7.2's list with: promotion stability and demotion rate. If hex wins, it wins as a COMPUTE topology — never retroactively as an explanation of the 96-bit register. + +### 7.5 Calibration-before-demotion — the instrument must qualify before it may judge (operator-hardened, 2026-08-25) + +> **A corpus may falsify a crystal only after it has falsified the +> hypothesis that it is an incompetent instrument.** + +**⊘ The defect this corrects was in MY first draft of the positive +control, and the operator caught it.** The draft rule — "a corpus may +only demote if the 25 established macros do NOT fall into its null +range" — contains an immortality loop: the very macros a corpus might +legitimately demote also decide whether the corpus is allowed to demote +at all. `25 fall into the null → corpus ruled bad → the 25 may not be +demoted.` A self-immunizing vocabulary is the anti-eigenvalue failure +wearing a lab coat. Recorded rather than silently fixed. + +**The rule:** corpus qualification must be INDEPENDENT of the candidate +it judges. + +``` +CORPUS QUALIFICATION (all six, before any demotion verdict) + +1. sufficient chains / effective sample size +2. no degenerate opcode monoculture +3. the null model's variance/range has not collapsed +4. known POSITIVE controls separate from the null (known signal → seen) +5. known NEGATIVE controls stay in/near the null (known noise → not seen) +6. the candidate under evaluation takes NO part + in its own corpus qualification +``` + +Point 5 is load-bearing: a positive control alone cannot catch a corpus +that makes EVERYTHING look strong. The instrument must demonstrate both +halves — `known signal → erkannt, known noise → nicht erkannt` — before +it may rule on an unknown. + +**Cheap immediate form — leave-one-out calibration:** + +``` +demote M17? + qualify the corpus using {established macros} \ {M17} + frozen negative controls + only then evaluate M17 +``` + +M17 can now actually die. + +**Target form — a frozen calibration panel, NOT identical to the System +palette:** + +``` +CAL+ a few provably portable R2IL macros (frozen at panel creation) +CAL− shuffled / corrupted / contract-invalid macros + (seed material exists: B4's corruption demo already constructs a + contract-invalid macro and proves the check catches it) + +held-out corpus → CAL+ separates? → CAL− stays null? → corpus is +COMPETENT → may contribute demotion evidence +``` + +**And even a qualified failure is EVIDENCE, not a verdict** — the §7.4 +hysteresis applied to demotion: + +``` +qualified failure → demotion evidence accumulates + → crosses the measured lower threshold + → System → Learned +``` + +Otherwise "never demote" is merely traded for "one strange Rust binary +in a bad mood empties the library." The accumulation threshold is a +POLICY PIN, measured when the first real demotion case exists. + +This is the same philosophy as every falsifier in this workspace, one +level up: not only the specimen must be falsifiable — **the instrument +must prove it can currently measure.** (Precedent, small scale: the +shuffle runs of 2026-08-25 asserted their multiset/marginal preservation +on every draw — each null was itself checked before its verdict counted.) From 3bf1ef59b1052f8f7e4efab68dd4c0293510ea5f Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 25 Aug 2026 16:38:55 +0000 Subject: [PATCH 5/6] =?UTF-8?q?plan(r2il-contract):=207.6=20=E2=80=94=20on?= =?UTF-8?q?tology=20morphogenesis=20as=20the=20probe's=20real=20task?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Supersedes 7.2's cheap first version as the primary form of PROBE-GREY-TOPOLOGY-AB: fold intra-ontology structure only, with the operator's leakage fence extended past explicit crosswalks to any derived feature that betrays them; score emergent basins against the withheld curated mappings; four nulls including the is_a-depth- preserving leaf permutation that kills common-root attraction; the 7.5 instrument qualification applied to the oracle itself so a broken oracle gets no veto; and a two-phase sequencing where novel basins become admissible only after known withheld relationships are recovered, entering as Explore-status crystals rather than findings. The data half stays in the consumer repo's own ledger. --- .../r2il-machine-semantic-contract-v1.md | 62 +++++++++++++++++++ 1 file changed, 62 insertions(+) diff --git a/.claude/plans/r2il-machine-semantic-contract-v1.md b/.claude/plans/r2il-machine-semantic-contract-v1.md index 038a3633d..9a32d88d1 100644 --- a/.claude/plans/r2il-machine-semantic-contract-v1.md +++ b/.claude/plans/r2il-machine-semantic-contract-v1.md @@ -415,3 +415,65 @@ level up: not only the specimen must be falsifiable — **the instrument must prove it can currently measure.** (Precedent, small scale: the shuffle runs of 2026-08-25 asserted their multiset/marginal preservation on every draw — each null was itself checked before its verdict counted.) + +### 7.6 §7.2's probe gets its real task: ONTOLOGY-MORPHOGENESIS (operator-approved + leakage-hardened, 2026-08-25) + +Supersedes the "cheap first version" in §7.2 (the 33-macro co-occurrence +graph) as the PRIMARY form of `PROBE-GREY-TOPOLOGY-AB`: independently +curated open ontologies (MONDO / HPO / GO / ChEBI — different projections +of one biological reality) are a far stronger oracle than internal +machine patterns held against each other. The baked corpus already +exists on the consumer side (the open-ontology SoA bake with +`(classid, identity)` lanes and `part_of:is_a` rails); the data half of +this probe lives in that repo's own ledger, only the topology half here. + +``` +PROBE-GREY-TOPOLOGY-AB / ONTOLOGY-MORPHOGENESIS (unbuilt, gated) + +INPUT: + exclusively intra-ontology structure: + is_a, part_of, roles, evidence, causal edges. + LEAKAGE FENCE (operator-added, load-bearing): not only explicit + crosswalks are withheld — DERIVED features that already betray them + are banned from the fold rule too: no xref-derived labels, no + pre-normalized shared IDs, no bridge features that carry the answer. + +WITHHELD ORACLE: + the known cross-ontology mappings (xrefs, curated anatomy mappings, + curated axis crosswalks). Hidden during folding; scored against after. + +COMPARE (same data, same fold rule, same budgets): + A = the substrate's native Morton/Cartesian neighbourhood + B = the experimental local topology (hex, if it applies) + +NULLS (all four; each must be qualified per §7.5 before it may veto): + degree-preserving rewiring (kills hub bias) + relation-type shuffle + label/identity shuffle + is_a-DEPTH-PRESERVING leaf permutation (kills common-root attraction — + without this, "everything near the root looks alike" gets celebrated + as morphogenesis) + +GATE: + A fold is interesting only if it survives destruction of the semantics + that supposedly caused it AND recovers withheld mappings better than + every qualified null/control. +``` + +**§7.5 applies to the oracle itself:** a null model or held-out set may +demote a fold only after proving it can separate known withheld +relationships as its positive control. A broken oracle gets no veto. + +**Sequencing, deliberately two-phase:** first the probe must RECOVER +what curators independently already knew, despite it being hidden — +self-supervised structure discovery with uncontaminated ground truth. +Only THEN do the novel basins (folds with no existing crosswalk) become +admissible as candidates — and they enter as Explore-status crystals +under §7.4/§7.5, never as findings. The ontologies themselves stay +separately true throughout: the geometry learns PROXIMITY, white matter +keeps the explicit edges, provenance, and truth. + +**Both outcomes pay:** if B/hex does not beat A on withheld recovery, +false resonance, and promotion/demotion stability, hex is dead and the +crystal stays square — a real answer. If it wins, the honeycomb has its +first empirical lease. From 35f10ae3a869ac1966b54461d7a805ee7b19f8dc Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 25 Aug 2026 16:39:14 +0000 Subject: [PATCH 6/6] =?UTF-8?q?plan(r2il-contract):=207.6=20addendum=20?= =?UTF-8?q?=E2=80=94=20the=20responsibility=20boundary,=20verbatim?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The plan defines the experiment's form and falsifiers, nothing else; corpus, bake state, withheld mapping lists and receipts stay in the consumer repo's own ledger. Stated as a tripwire for future edits: any ingestion or crosswalk-list text appearing here means the cut failed. --- .claude/plans/r2il-machine-semantic-contract-v1.md | 14 ++++++++++++++ 1 file changed, 14 insertions(+) diff --git a/.claude/plans/r2il-machine-semantic-contract-v1.md b/.claude/plans/r2il-machine-semantic-contract-v1.md index 9a32d88d1..6b30e6057 100644 --- a/.claude/plans/r2il-machine-semantic-contract-v1.md +++ b/.claude/plans/r2il-machine-semantic-contract-v1.md @@ -477,3 +477,17 @@ keeps the explicit edges, provenance, and truth. false resonance, and promotion/demotion stability, hex is dead and the crystal stays square — a real answer. If it wins, the honeycomb has its first empirical lease. + +**Responsibility boundary (operator, same day — the cut that keeps this +plan from owning MedCare):** + +``` +lance-graph plan: defines the topology experiment + falsifiers. NOTHING ELSE. +MedCare-rs: owns corpus, bake state, withheld mapping lists, receipts + (its own ledger, RAIL_OFFENE_POSTEN). +``` + +§7.6 defines the FORM of the experiment. The moment a future edit here +starts specifying ontology ingestion, bake artifacts, or concrete +crosswalk lists, the clean cut has failed — move that text to the +consumer repo and leave a pointer.