Skip to content

Accept-workflow and honest reporting: adopt what Playwright, VCR and SimpleCov already proved #259

Description

@pftg

Evidence from a lightning-demo teardown (.ai/lightning-demos.md) — every mechanism below was executed on a real machine, not read about. Transcripts are in that file.

Two failure moments drove this, both measured on real customers:

  • M1 — the miss. A green bar on a deliberately broken page.
  • M2 — the accept. The most frequent daily action, and we have no verb for itgrep -rni "def accept|approve|update_baseline" lib/ returns nothing, there are no rake tasks and no exe/.

#255 and #257 fixed the silence (we now warn loudly and the docs are true). This issue is about the shape.


1. Write the baseline on a miss — and fail anyway

Playwright's default missing mode, verified live:

Error: A snapshot doesn't exist at /.../homepage-darwin.png, writing actual.
  1 failed

toMatchSnapshot.ts:211 returns a literal false — the test fails while the file is written. Second run: 1 passed.

The user never has to learn how to record; recording happened. They get a red bar plus a file on disk, which is the honest signal: I could not verify this, here is what I saw, look at it and re-run.

This needs a decision, because it changes a default. Our fail_if_new defaults to ENV["CI"], so locally a new screenshot passes. Playwright's position is that a snapshot with no baseline has not been checked, so it is red. I think Playwright is right and both our customer personas were burned exactly here — but it makes "first run always passes" (README Quick Start) no longer true, and every existing user adding a test would see one new red bar. Wants 4-eyes before implementing.

The trap, stated by the researcher: writing-on-missing is safe because the run fails. Adopt the write without keeping the failure and we ship M1 as a feature.

2. rake snap_diff:accept — re-run, don't copy

Three independent tools converged on this shape: Playwright --last-failed -u, BackstopJS backstop approve (45 lines, promotes only failed_diff_*), Lost Pixel's update mode. RSpec's example_status_persistence_file_path is the Ruby-native precedent our users already know.

The property worth stealing: it re-runs the test rather than copying the previous run's artifact, so you cannot promote a baseline that no longer reproduces. Add FILTER= (BackstopJS has --filter) and narrate before acting.

The local equivalent of Percy's "approve" is git add of a file the tool just regenerated and re-verified for you. The verb is missing; the concept already fits.

3. Rewrite both messages (VCR + WebMock + SimpleCov)

VCR's unhandled-request error, produced live, has four properties we have none of: it restates the effective configuration back at you (:record => :none, :match_requests_on), enumerates every escape route including "if you're surprised, here's how to debug", and uses versioned doc links (?v=6-4-0) that cannot point at docs for an API you do not have.

WebMock goes further and emits the fix, pre-filled from this run's real values, rather than describing it.

Design rule, and it is why RECORD_SCREENSHOTS rotted for years: never print a command in an error message that is not generated from live state.

SimpleCov's threshold failure earns every one of its four lines — what it is, what was required, where to look first, and a distinct exit code 2 for "the tool ran fine, your content failed the gate".

4. Labelled, relative, ordered paths

Expected: tests/vis.spec.js-snapshots/homepage-darwin.png
Received: test-results/.../homepage-actual.png
Diff:     test-results/.../homepage-diff.png

Ours prints four unlabelled absolute paths and the reader guesses. Playwright omits Diff: when there is none — the block is honest about what exists.

5. A third state in the summary line

Argos names the no-baseline state orphan — not a pass, not a failure. Ours should count it:

14 screenshots verified, 0 changed, 1 new (not verified)

You cannot read that and believe the page was checked. Highest value-per-byte item here.


The anti-lesson, from the market leader

Percy exits 0 when it silently does nothing:

$ npx @percy/cli exec -- echo hi
[percy] Skipping visual tests
[percy] Error: Missing Percy token
[percy] Command "echo hi" exited with status: 0
$ echo $?
0

It prints the word Error, says it is skipping every visual test, and exits successfully. A CI job that loses its token goes green forever. That is our M1 in its purest form, shipped by the biggest name in the category.

Contrast Chromatic, where every route to a green bar without a real comparison has its own named, branch-scoped flag — --auto-accept-changes, --exit-zero-on-changes, --skip. You cannot land there by accident, and a reviewer reading CI config sees it.

Rule: misconfiguration must be a non-zero exit.


Cost

Items 1–5 are one config key, one rake task, two rewritten methods and a counter. No subsystem. Target 2.1; items 4 and 5 are cheap enough to consider for 2.0 if the release slips.

Deliberately not stealing: Percy/Chromatic/Argos dashboards (need a service), TurboSnap dependency tracing (four flags to configure one feature), jest-image-snapshot's terminal image embedding (two options and a terminal allow-list for something a labelled path does for everyone), a .snap_diff_todo.yml (our accepted state is already PNGs in git — more reviewable than a YAML exclusion list).

Also worth a follow-up: stale-baseline detection (jest-image-snapshot's OutdatedSnapshotReporter, ~20 lines). First version must report only, never delete — obsolescence is only knowable from a full run, and deleting after a filtered run destroys baselines.

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions