From b9b5ecb059a6c55ba88e18a4ea701562e499ca3c Mon Sep 17 00:00:00 2001 From: jgoneit Date: Sat, 29 Aug 2026 08:12:01 +0900 Subject: [PATCH] docs(eval): define terminal observation boundary Why: - Assistant turn completion could be mistaken for a terminal task outcome, causing premature or missing observations across resumable tasks. What: - Define nonterminal waiting and recoverable Ward or Seal stop states. - Narrow completed, failed, and Host-supplied abandoned outcomes. - Require one observation only after the eventual terminal outcome. - Document the Ward state-root preflight and aborted-run restart boundary. Validation: - gofmt -l . - go vet ./... - go test ./... - go test -race ./... - go build ./cmd/evalctl - Ward exact-root permission smoke and multi-turn Host E2E --- README.md | 62 +++++++++++++++++++++++++++++++++++++++++++++++++------ 1 file changed, 56 insertions(+), 6 deletions(-) diff --git a/README.md b/README.md index b89c475..eec199c 100644 --- a/README.md +++ b/README.md @@ -84,16 +84,61 @@ Legacy Eval v1/v2 state is neither read nor migrated. ## Host experiment policy The experiment population is the next 20 eligible root Codex tasks after the -managed host policy becomes active. Codex loads global `~/.codex/AGENTS.md` +managed host policy becomes active. The Host launches every candidate with the +preflighted `ward` permission profile. Launching a candidate with another +profile is a provisioning failure that aborts the experiment run; it is not a +reason to exclude that task afterward. Codex loads global `~/.codex/AGENTS.md` guidance at session start, so only fresh tasks are eligible after installation. See the [OpenAI AGENTS.md documentation](https://learn.chatgpt.com/docs/agent-configuration/agents-md). -Eligible terminal outcomes are `completed`, `failed`, and `abandoned`. Eval -development, pure Q&A, and subagent child tasks are excluded in advance. Ward -or Seal usage and the quality of the result never exclude a task afterward. +Before activation, the Host adds only the Eval state directory +`$XDG_STATE_HOME/jgoneit/eval-experiment` (or the platform fallback shown +above) as an explicit writable workspace root for the selected permission +profile, creates its private directories with mode `0700`, ensures any files it +creates use mode `0600` on Darwin/Linux, and verifies writes against a separate +preflight state root. For a managed permission profile, this workspace-root +rule is authoritative and must not be combined with legacy +`sandbox_workspace_write` settings. See the [OpenAI Permissions documentation](https://learn.chatgpt.com/docs/permissions). + +A preflight always passes a disposable absolute child path beneath that granted +Eval state directory through `--state-root` and confirms that its journal path +differs from the production journal before it runs. It must not create or +modify the production journal. If state provisioning or lifecycle calibration +fails, the Host marks that experiment run aborted and restarts with a new +activation time; it does not silently reuse the failed run or migrate its +state. If an aborted run already contains rows, the Host quarantines that exact +journal subtree before starting the replacement run rather than appending new +samples to it. + +The end of an assistant turn is not by itself a terminal task outcome. The +same primary objective remains nonterminal only while an identified, in-scope +continuation can resume it after user input, approval, configuration, or +permission becomes available. A recoverable Ward or Seal stop, clarification +request, or resumable handoff is therefore nonterminal and must not be mapped +to `failed` or `abandoned`. + +Eligible terminal outcomes have these narrower meanings: + +- `completed`: the requested objective is achieved and no required work remains; +- `failed`: the objective cannot be completed within the current scope and + granted authority, no identified in-scope continuation remains pending, and + the task is being closed; +- `abandoned`: the Host explicitly supplies the outcome; the Agent never infers + abandonment from user silence. + +Without an explicit Host terminal callback, an abandoned task contributes only +to the Host denominator and missing-attempt count; it cannot produce an Agent +observation. + +Eval development, pure Q&A, and subagent child tasks are excluded in advance. +Ward or Seal usage and the quality of the result never exclude a task afterward. For an eligible task, the primary Agent makes one silent `evalctl observe` -attempt after the task outcome is fixed. It does not: +attempt only after the eventual terminal outcome is fixed. It makes no attempt +during a nonterminal turn; if the same task resumes, it preserves eligibility +for the single terminal attempt. The attempt pipes one complete JSON object to +standard input in the same shell command; bare `evalctl observe` is invalid and +must never be called. It does not: - run Ward or Seal to discover a metric or version; - ask a question or request approval; @@ -102,7 +147,12 @@ attempt after the task outcome is fixed. It does not: The Host task history supplies the population denominator and missing-attempt count. The journal intentionally does not store skipped rows or task identity. -The managed host block is removed after the twentieth eligible task. +The managed host block is removed after the twentieth eligible task. If the +same objective reopens after a terminal attempt, the Agent makes no second +attempt. The Host treats that reopen as a lifecycle-calibration failure, aborts +the experiment run, and restarts with a new activation time rather than +continuing with the stale row. This minimal experiment does not correct or +supersede recorded rows. ## Decision boundary