Skip to content

Say what justifies a release, and what does not - #81

Merged
leggetter merged 1 commit into
mainfrom
release-cadence-policy
Sep 21, 2026
Merged

leggetter merged 1 commit into
mainfrom
release-cadence-policy

Conversation

@leggetter

Copy link
Copy Markdown
Collaborator

From the question "are we due a release?" — which the convention could not answer.

It said "cut one when there is a measured change to report, not on a schedule", and left the three cases anyone actually faces open: time passing, the benchmark changing, the product changing.

Three triggers, any one sufficient

  1. The product moved — something shipped that this benchmark found, or a finding worth publishing came out of a run.
  2. The instrument changed what it measures, and a full run has been measured under it — a scenario, a scorer, the base prompt, the CLI pin. Until then the page shows numbers measured under something the notes do not describe.
  3. The snapshot has gone stale — roughly six weeks, sooner if the instrument moved under it. This one has a deadline rather than a preference: transcripts expire at ninety days (Publish transcripts with releases so a published cell can be checked #21), so a snapshot nobody released is evidence nobody can check later.

One explicit non-trigger

A run whose numbers moved. Movement is the default here. The weak model's delta read −1 on 1 September and +4 on 14 September with nothing changed between them; eleven cells flipped and two moved in both directions across the arms of the same pair (#2). Releasing on movement publishes variance and spends the changelog on non-events.

Package the monthly, not a weekly

A weekly covers four experiments, so its snapshot carries the frontier -no-skills arms forward from whenever they last ran. Publishable, but a release whose notes say "one run" would be wrong — which is the property v0.4.0 was cut for.

🤖 Generated with Claude Code

https://claude.ai/code/session_01MQzUoMAwEBJWpEGVvVzSjK

The convention said "cut one when there is a measured change to report" and
left the three cases anyone actually faces unanswered: time passing, the
benchmark changing, the product changing.

Three triggers, any one sufficient. The product moved — something shipped or a
finding came out of a run. The instrument changed what it measures and a full
run has been measured under it, so the numbers and the notes describing them
agree. Or the snapshot has gone stale, which has a deadline attached:
transcripts expire at ninety days, so a snapshot nobody released is evidence
nobody can check later.

And one explicit non-trigger, because it is the tempting one. A run whose
numbers moved is not a reason. The weak model's skills delta read -1 on
1 September and +4 on 14 September with nothing changed between them; eleven
cells flipped and two of them moved in both directions across the arms of the
same pair. Releasing on movement publishes variance and spends the changelog
on non-events.

Also records that the monthly matrix is the better thing to package: a weekly
covers four experiments, so its snapshot carries the frontier -no-skills arms
forward and a release calling it "one run" would be wrong.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MQzUoMAwEBJWpEGVvVzSjK
@leggetter
leggetter merged commit d8f1c0a into main Sep 21, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant