Remove evidence-certification bookkeeping; keep the benchmark harness - #126
Merged
Merged
Conversation
This was referenced Sep 17, 2026
Deletes the layer that certified past benchmark runs: sealed archives, admission registries, sha256-pinned result manifests, the generated conformance ledger, and the tests that assert those digests exist. The harness that produces measurements stays: datasets, differential configs, upstream pins, benchmark mix tasks, and the tests of harness behaviour. Tracked tree drops from 1599 files / 87.9 MB to 1185 files / 14.5 MB.
Keeps every test that exercises a surviving task, campaign runner, scorer or differential against a local fixture or the pinned upstream, together with the configs, datasets and upstream pins those readers need. Only assertion blocks that call a deleted validator or registry, or that read a deleted archive, are removed.
…fact-path assertion
deepfates
force-pushed
the
claude/cut-evidence-bookkeeping
branch
from
September 17, 2026 19:33
674c308 to
22cdce0
Compare
This was referenced Sep 17, 2026
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What a reader of the public repo gains
Before this change, most of the tracked bytes were a record that certain benchmark
runs happened: sealed
.tar.zstarchives, content-addressed "admitted evidence"files, registries binding claims to those files, and tests whose only assertion was
that a digest still matched. None of it let an outside developer run anything.
What is left is everything a stranger can run. The dataset fetchers and loaders, the
scorers, the campaign runners, the benchmark mix tasks, the differential harness
against pinned DSPy and GEPA, the upstream pins it reads, and every test that
exercises a task, a campaign runner, a scorer or a differential against a local
fixture or the pinned upstream.
mix differential.checkruns the whole:dspy_paritylane unchanged.Clone size: 1599 tracked files / 87.9 MB becomes 1249 files / 15.3 MB.
Deleted, by category
evidence/sealed archivesbenchmarks/results, runs, admitted evidence, locks, claim and reproduction registriesexamples/matched-experiment output dumpstest/certification testsdocs/maintainer registry docs and the generated conformance ledgerlib/mix/tasks/registry and coverage-ledger tasksbench/registry and admission modulesscripts/with no runnable calleridentity/.github/workflows/evidence.yml350 files, 72.5 MB.
What was kept, and why
benchmarks/data/,benchmarks/config/,benchmarks/authorities.json(whole, includingits
familiessection — five differential Python scripts and four differential mix tasksread it),
benchmarks/authority_sources/,benchmarks/upstream/gepa-corrected/(thepatch
scripts/setup_corrected_gepa_comparator.shapplies in the CI differential job),and every requirements lock the setup scripts install.
Every benchmark mix task that runs something is kept, including
imp.benchmark.parity.aggregate,imp.benchmark.gepa_replication,imp.benchmark.gepa_merge,imp.benchmark.optimize_anything,imp.benchmark.support_ticket_lift,imp.benchmark.provider_training,imp.benchmark.instruction_optimizer_campaign, the two OpenRouter free-route canaries,and
imp_acp.demo_mcp_http_server. Missing documentation for a runnable task is adocumentation gap, not a reason to delete the task.
Mix tasks deleted
imp.evidence.admit,imp.evidence_authorities,imp.reproductions,imp.research_portfolio,imp.upstream_fidelity— registry validators anddocumentation generators for ledgers that no longer exist.
imp.benchmark.live_matrix— a coverage ledger over stored campaign artifacts, keyed onthe deleted
benchmarks/model_availability.json; it calls no provider and measuresnothing.
imp.benchmark.hotpotqa_analysis— a post-hoc classifier over stored parityartifacts.
Bench modules deleted
Imp.EvidenceAuthorities,Imp.ReproductionRegistry,Imp.ResearchPortfolio,Imp.UpstreamFidelity,Imp.BenchmarkTruth.EvidenceAdmission,Imp.BenchmarkTruth.ReproductionArtifactValidator. Verified unreachable from everysurviving lib or mix task, bench module, script and test.
Imp.BenchmarkTruth.Pathslosesadmitted_root/0andadmitted/1.Imp.BenchmarkTruth.OptimizeAnything.Campaignreadsbenchmarks/authorities.jsondirectly instead of through the deleted loader.Imp.BenchmarkTruth.SupportTicketLiftCampaignno longer digests two deletedpredecessor result files while validating a preregistration manifest.
scripts/dspy_optimizer_public_workflow_gate.pydefines its three-lineOperationalSafetyAbortlocally instead of importing it from a deleted example dump —that class is what the gate measures.
Tests: one line each
Deleted whole — bookkeeping (its only assertions are against a deleted validator,
registry or archive digest):
claims_inventorybenchmarks/claims.jsonagainst maintainer docsevidence_admissionevidence_authoritiesevidence_provenanceorigin/mainand is cited by a ledgerreproduction_registrybenchmarks/reproductions.jsonresearch_portfoliobenchmarks/research_portfolio.jsonupstream_fidelitytutorial_ticket_routing_artifactDeleted whole — behaviour lost because its subject is gone:
experiment_reference_graphclaims.json, and a sha256 migration ledger;scripts/experiment_reference_graph.exsgoes with itmatched_gepa_mipro_ifbenchexamples/matched_gepa_mipro_ifbenchdumpsmatched_ifbench_stop_accountingstop_accounting.exsfrom two deleted example dirsmatched_instruction_family_ifbench_designmatched_instruction_optimizers_trec_examplecase_study_trec_recomputationexamples/and restores this test therelocal_optimize_anything_retry_policy_three_seed_evidenceRestored after the first review — behaviour, kept with only the deleted-validator
assertions removed:
instruction_optimizer_contract_artifactimp.benchmark.instruction_optimizer_contractagainst pinned DSPy; 33 structural rows and the cross-runtime bootstrap fixtureReproductionArtifactValidatorcalls and the tamper blockgepa_contract_artifactGepaContract.compareagainst pinned GEPA v0.1.4; 15 required casesadmittedenvelope and its validator callssupport_ticket_lift_campaign(v2, v3, v3 full runner)SupportTicketLiftCampaign.runthrough a local HTTP fixture for 48 requests and checks the response ledger; the fourbenchmarks/configpreregistration files are restoredoptimize_anything_campaignoptimize_anything_artifactoptimize_anything_agent_config,optimize_anything_code_artifact,optimize_anything_scheduling_heuristic,optimize_anything_v1_4_readiness,circle_v1_4_cold_plangepa_replication_artifact,gepa_merge_task,gepa_campaignlangprobe_heart_disease_product_fitmusique_ans_product_fit,musique_ans_mipro_currenthover_papillon_calibration_pilot,hover_gepa_no_merge_plangepa_study_planbenchmark_pathsPaths.admittedassertionscopro_isolation_artifactauthority_inventorybenchmarks/authorities.jsoninternally valid — the ledger the differential lane readsclaims.jsonand the deleted conformance docTrimmed, subject partly gone:
benchmark_truthlive_matrixorhotpotqa_analysis, whose tasks are deleted. Nothing else was removed; the parity aggregator, fetcher, runner, integrity and scorer tests are untouched.local_mlx_campaignconfidence_calibration_benchmarkbfcl_adapted_artifactupstream_authority_registryUpstreamFidelityassertions removed; the registry fail-closed tests stay and a pin-resolution test replaces the deleted one.documentation_contractoperations_stresspackage_contractConsistency
mix.exsloses only the aliases whose tasks are gone (upstream_fidelity.check,reproduction.check,research.portfolio.check,benchmark.live_matrix,benchmark.hotpotqa_analysis) and theirpreferred_envsentries.benchstays inelixirc_pathsfor dev and test..github/workflows/ci.ymlneeds no change: its jobsrun mix aliases, and every path in the
paths-filterlists still exists..dialyzer_ignore.exsdrops the entry for the deleted task and re-pins one line thatmoved.
.gitignoreno longer ignores tracked files.test/test_helper.exskeeps the:evidence_infrastructureexclusion — 40 tagged files remain.CONTRIBUTING.md,decisions.mdanddocs/EVIDENCE.mdlose their pointers to deleted paths.Verification
MIX_ENV=test mix compile --warnings-as-errors: clean.mix format --check-formatted: clean.mix quality.check(credo,mix deps.audit,mix hex.audit— the exact CI commands): exit 0.mix testonmain:54 doctests, 9 properties, 2998 tests, 0 failures, 11 skipped (171 excluded)mix teston this branch:54 doctests, 9 properties, 2884 tests, 0 failures, 10 skipped (146 excluded)mix dialyzer:Total errors: 150, Skipped: 150, Unnecessary Skips: 0—done (passed successfully).:evidence_infrastructuresuites outside the default run were exercisedseparately against a provisioned pinned environment: the COPRO isolation
differential and artifact, the hover pilots, the GEPA contract, the BFCL scorer and
the authority inventory all pass.
check,stranger.check,classify changes,package.check,integration.check,protocol.check,quality.check,dialyzer.check,differential.check(11m01s),docs.check.Not in this change
docs/CASE_STUDY_TREC.mdanddocs/internal/still reference deleted paths; thebenchmark-surface branch owns both.
Everything deleted here exists as a tarball on the owner's machine.