feat(pr-workflow): add /mms-debug orchestrator - #98
Conversation
`pr-validate` is handed a claim and looks for the observation that would falsify it. Debugging starts from a symptom and has to generate the hypothesis first, which is where the expensive failure lives: the theory is yours, nobody else is positioned to challenge it, and confirmation is cheap. Routes the symptom to the engine that owns its defect class rather than reimplementing any investigation, so both orchestrators share one set of engines. Carries `pr-validate`'s trust gates, which bind harder here — a weak instrument in review yields a claim someone challenges, in debugging it yields a theory nobody checks.
/mms-debug orchestrator
Context budgetWhat this PR costs an agent, measured from an install rather than read from the diff. Three tiers, and only the first is unavoidable.
Frontmatter is the only tier paid unconditionally — every agent loads it on every run once the skill is installed, used or not, because it is what the agent reads to decide relevance. The 28 skills across the eleven open skill PRs sit at a median of ~1,716 tokens selected and ~1,860 with references followed. All are within the 1,536-character description budget. Selected is paid only when the agent picks the skill. + refs & knowledge is the ceiling if every bundled reference is then read; it is a worst case, not an expectation. Method
These figures are pinned to the commit above and drift on every push; #96 tracks automating them. |
The installer emits `mms-debug`; the description advertised `/debug`.
The symptom table and the description both named `react-render-proof`. No skill of that name exists on `main` or in any open pull request; the render engine is `react-render-delta`, added by #43 and carried by #84 and #108, and installed locally as `mms-react-render-delta`. The other six engines this skill routes to all resolve to skills in open pull requests, so this was the only wrong name rather than one of three. Checked against the repository rather than against an install: none of the seven is on `main` yet, which is a merge-ordering fact and not a defect here. Nothing validates this today. #103, which would have checked cross-skill references, was closed as superseded by #87, and #87 is not merged — so this name would have shipped unflagged.
Overview
Adds one skill,
debug, atdomains/pr-workflow/skills/debug/skill.md. Its instructions are a routing table: hand it a symptom, it routes to the skill that owns that defect class —memory-leak,race-condition-repro,react-render-delta,sentry-grafana-correlation,extension-errors-debugging,tsc-blindspots,supply-chain-audit. It reimplements none of them; the routing decision is the addition, not the investigation.Motivation
Those seven are invocable by anyone who knows which they need;
debugis for when you have a symptom and do not. It isevidence's symptom-first sibling (#84), whose trust gates it carries:evidencestarts from a claim and looks for what would falsify it;debugmust generate the hypothesis first. Killed hypotheses are part of the output.debugalso leans ondistinguishing-observationandobservability-gap, from #106.Showcase
node .github/scripts/lint-skill-entry.mjs domains/pr-workflow/skills/debug/skill.md— 0 errors, 0 warnings