Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
35 commits
Select commit Hold shift + click to select a range
044cb93
feat: let an analysis cite from an explicitly versioned corpus
DavidHLP Sep 28, 2026
2b2abf1
fix: call the corpus override a test seam, not an authorization path
DavidHLP Sep 28, 2026
a567718
fix: gate an injected corpus with a manifest and filter status before…
DavidHLP Sep 28, 2026
b380dd9
fix: apply the source-size cap to a supplied corpus
DavidHLP Sep 28, 2026
eeffb61
fix: validate a supplied manifest through load_manifest, not just its…
DavidHLP Sep 28, 2026
368e144
feat: give the acceptance entry point a manifest-gated corpus path
DavidHLP Sep 28, 2026
cbf5575
test: prove the acceptance entry point takes authorised material end …
DavidHLP Sep 28, 2026
8bc1928
Merge origin/main into the acceptance-corpus branch
DavidHLP Sep 28, 2026
c223030
fix: harden the corpus override and document its operator contract
DavidHLP Sep 28, 2026
fa73caa
docs: name every corpus-override failure and say where the key comes …
DavidHLP Sep 28, 2026
9b05c4c
fix: close the remaining override gaps from the second review pass
DavidHLP Sep 28, 2026
d860ec9
docs: match the override reason list to the second review pass
DavidHLP Sep 28, 2026
4ad4034
fix: reject copied fragments and read entries through one no-follow d…
DavidHLP Sep 28, 2026
433e49e
test: make the no-follow regression indistinguishable from a valid read
DavidHLP Sep 28, 2026
ae380ae
fix: anchor the override to one root and carry one validated snapshot
DavidHLP Sep 28, 2026
28064e9
fix: parse the manifest from one captured read
DavidHLP Sep 28, 2026
183321b
fix: bind the override at preflight and refuse an unreachable threshold
DavidHLP Sep 28, 2026
42be9f3
fix: send the citation judgement inside the adapter answer envelope
DavidHLP Sep 29, 2026
dfd157f
test: pin the answer-level columns the development split is missing
DavidHLP Sep 29, 2026
6467756
feat: add the answer-level evaluator for the development split
DavidHLP Sep 29, 2026
931e29a
test: make the empty-retrieval fixture actually retrieve nothing
DavidHLP Sep 29, 2026
2fd1080
feat: add the development-split answer evaluation entry point
DavidHLP Sep 29, 2026
7313ea4
fix: retry a stalled answer-evaluation call instead of the whole batch
DavidHLP Sep 29, 2026
46371f5
docs: document the development-split answer evaluation entry point
DavidHLP Sep 29, 2026
fae735f
fix: classify observed behaviour from the answer text, not its own label
DavidHLP Sep 29, 2026
051d12e
fix: raise the answer-evaluation output cap and name truncation failures
DavidHLP Sep 29, 2026
caea09c
fix(agent): bind evals to validated snapshots
DavidHLP Sep 30, 2026
a77b4d1
fix(agent): harden answer evaluation failures
DavidHLP Oct 1, 2026
fbc2a5e
fix(agent): validate answer citations and corpus policy
DavidHLP Oct 1, 2026
27b2f81
fix: harden evaluation artifact and judge input boundaries
DavidHLP Oct 1, 2026
9bce536
fix: anchor corpus ancestors and preserve aborted evaluation usage
DavidHLP Oct 1, 2026
b635757
fix: isolate untrusted citation judge fields
DavidHLP Oct 1, 2026
335f981
fix: reject corpus FIFOs without blocking preflight
DavidHLP Oct 1, 2026
4ef57f7
fix: normalize corpus duplicates and protect input paths
DavidHLP Oct 1, 2026
73375e9
fix: anchor artifact lifecycle and normalize manifest depth failures
DavidHLP Oct 1, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
33 changes: 32 additions & 1 deletion docs/DEVELOPMENT.md
Original file line number Diff line number Diff line change
Expand Up @@ -69,6 +69,24 @@ uv run pytest -q

真实 UltiCode HTTP / 模型 e2e 仍是显式 opt-in;`e2e_sourced_analysis.py` 使用 agent-authored synthetic Markdown corpus(不是提交、DTO 或用户授权材料),分析输入则是 authenticated user 的 validated read-only submission projection,且不调用真实模型。`e2e_sourced_analysis_model.py` / `e2e_model_qa.py` 是真实模型入口,调用形如 `DEEPSEEK_API_KEY=... DEEPSEEK_MODEL=<model> uv run python e2e_sourced_analysis_model.py`,仍需显式提供现有环境和 `DEEPSEEK_API_KEY` 与 `DEEPSEEK_MODEL`;不得把本地开发账号密码、Cookie、源码、检索文本或模型回答写入日志。可执行题集和当前评估状态见 `services/agent/data/keyword_cases.json` 与对应 Linear 任务。

`e2e_answer_evaluation.py` evaluates generated answers on the development split only. Its answer
pass receives the case question and retrieved evidence, not expected/allowed/forbidden outcomes;
it must return explicit retrieved chunk IDs in `citations`, and judging checks only those citations.
Holdouts remain sealed. Export the existing `DEEPSEEK_API_KEY` before this opt-in call; the entry
requires `DEEPSEEK_MODEL`, budgets two **logical** passes per development case (answer then judge;
default ceiling 64 — a retried transport error or timeout is billed again and counted in the row's
`model_calls`); usage is emitted from the model session even when a later answer or judge
pass aborts, with unknown token totals labelled `unknown`. The runner treats the generated answer as one untrusted JSON string value and tells the
judge to ignore directives inside it (this is a prompt boundary, not proof of injection immunity), reserves the artifact destination before the first billed call and never overwrites
an existing one, snapshots the corpus and case file once before the calls so the artifact identifies
the material actually judged, and writes results under the user state directory without printing
answer text:

```bash
cd services/agent
ULTICODE_ANSWER_EVAL=1 DEEPSEEK_MODEL=<model> uv run python e2e_answer_evaluation.py
```

真实模型入口的调用形式(`DEEPSEEK_MODEL` 为必填,脚本不再继承任何默认模型标识。`DEEPSEEK_MAX_TOKENS` 默认值按入口不同:`e2e_sourced_analysis_model.py`(每次运行仅 1 次调用)默认 2000——实测 300 时该次决策被输出上限截断(`finish_reason=length`、`content_len=150`)而报 `ModelProtocolError`,2000 时同一 prompt 通过;`e2e_model_qa.py`(最多 8 次调用)保持 300,真实运行在 300 下即通过,不应无证据地抬高其每次上限):

```bash
Expand All @@ -94,7 +112,20 @@ uv run python e2e_citation_support_model.py

`DEEPSEEK_API_KEY` 只在真正进入判定阶段才需要:检索、完整性门禁与引用条数不足(`insufficient_citations`)都在**只读预检**里完成,因此没有凭据也能看到材料缺口。

它只判定**分析实际发出的引用**:每条引用一次模型调用,只问「片段是否支持结论」(`supports` / `derivable`);`exists` 始终取自确定性完整性门禁,不由模型决定。verdict 写盘后经 `load_verdicts` 读回再汇总,因此仍按重算 id 绑定到具体 claim/quote。少于 `ULTICODE_CITATION_REQUIRED_ROWS`(默认 3)报 `insufficient_citations`;未过**完整性门禁**的引用在**发起任何模型调用之前**就报 `citation_integrity_failed`;verdict 汇总阶段发现不支持的引用报 `citation_gate_failed`。这些都以至退出码 1 结束,不报「低分通过」。输出行标注 `judge=model` 与 `human_review=not_performed`,以便与将来的人工复核记录区分。verdict 默认写到**状态目录**(`$XDG_STATE_HOME` 为绝对路径时用它,否则 `~/.local/state`)下的 `ulticode/citation-verdicts-<随机>.json`:每次运行独立,且不落在 checkout 里,可用 `ULTICODE_CITATION_VERDICTS` 改路径;无论哪种,目标位置都会在**付费调用之前**被独占占位,不可写或已被占用即 `verdict_destination_unusable`;写盘同时生成 `<path>.meta.json` 附属文件,记录判定者(`judge=model`)、模型标签、语料、阈值与提交事实摘要,使 verdict 脱离本次会话仍可解释。密钥只从环境读取,不得写入日志或仓库;上面两条真实模型入口示例中的 `DEEPSEEK_API_KEY=...` 只是占位符,实际运行同样应先在环境中导出。
它只判定**分析实际发出的引用**:每条引用一次模型调用,只问「片段是否支持结论」(`supports` / `derivable`);`exists` 始终取自确定性完整性门禁,不由模型决定。verdict 写盘后经 `load_verdicts` 读回再汇总,因此仍按重算 id 绑定到具体 claim/quote。可选地用仓库外的语料替换默认样本(**两个变量都要或都不要**,否则 `FAIL reason=corpus_source_incomplete`,且不会回退到默认语料):

```bash
# DEEPSEEK_API_KEY 必须已由操作者导出或由密钥存储注入到环境,绝不出现在本命令行;
# 缺失时以 deepseek_api_key_required 结束。
ULTICODE_CITATION_SUPPORT=1 DEEPSEEK_MODEL=<model> \
ULTICODE_CITATION_CORPUS_DIR=/绝对路径/语料目录 \
ULTICODE_CITATION_CORPUS_MANIFEST=/绝对路径/manifest.json \
uv run python e2e_citation_support_model.py
```

声明一律以 manifest 为准(缺失/不可读/非法或嵌套过深 JSON/校验不过 → `corpus_manifest_unusable`),每条必须逐字等于源码里钉死的**材料类别**——`ACCEPTED_PERMISSION` / `ACCEPTED_SCOPE` / `ACCEPTED_SAMPLE_KIND` / `ACCEPTED_ACCESS_SCOPE`(授权材料时四个一起改)(`corpus_declaration_unsupported`);manifest 的祖先路径也逐段通过 no-follow 目录描述符打开,拒绝祖先符号链接;manifest 必须是以 `O_NOFOLLOW|O_NONBLOCK` 打开的普通文件,按最多 1 MiB 有界读取并确认 EOF(符号链接、FIFO、非普通文件及超限均报 `corpus_manifest_unusable`);manifest **只读一次**,其条目、对应文档与钉死类别作为同一份不可变快照贯穿检索、worksheet 与 verdict 元数据,运行中途替换文件不会让它们描述不同材料;根目录必须是真实目录而非符号链接(`corpus_root_unusable`),manifest 至少声明一条(`corpus_empty`);根路径从 `/`(绝对路径)或 cwd 描述符(相对路径)开始逐段以 `O_DIRECTORY|O_NOFOLLOW` 打开,祖先符号链接也被拒绝,条目相对该描述符打开(根与条目在检查与读取之间被换成链接都会被内核拒绝),按 manifest 自身 `source_path` 的**纯文件名**解析(绝对路径或外部路径不是实际打开的文件,报 `corpus_entry_path_not_relative`;引用里记录的 `source_path` 即该已验证文件名),一条对应一个文件,符号链接(`corpus_entry_escapes_root`)、缺失(`corpus_entry_missing`)、两个条目指向同一个**文件**(含同一 inode 的两个硬链接名,`corpus_entry_duplicate_source`;副本按规范化文本比较,Unicode NFC 规范等价、CRLF/LF 或无关尾随空白差异也算同一片段;NFC 仅用于去重,不改变原文、原始摘要与行号 → `corpus_entry_duplicate_content`;条目用 `O_NOFOLLOW|O_NONBLOCK` 单次打开,非普通文件先经 `fstat` 拒绝,FIFO 无 writer 也不会阻塞预检,检查与读取之间被换成符号链接也由内核拒绝——总之单片段不得被计成多条引用)、不可读/非 UTF-8/为空/超 `MAX_SOURCE_CHARS`/有界读取未到 EOF(前缀过检而后缀未读)(`corpus_entry_unusable`)都直接失败;空 manifest 列表单列为 `corpus_empty`(区别于解析不了的 `corpus_manifest_unusable`);`source_position` 由**原始文件**推导(首末非空物理行,仅 LF、CRLF、CR 算换行,前导空行也计入),与 manifest 不符即 `corpus_entry_position_mismatch`。不设这两个变量时行为与默认完全一致。证据行与 verdict 元数据里的 `corpus=` 取自钉死的 permission,因此非 synthetic 材料不会被标成 synthetic。契约细节见 `services/agent/README.md`。

少于 `ULTICODE_CITATION_REQUIRED_ROWS`(默认 3)报 `insufficient_citations`;阈值**高于检索上限**(`MAX_RESULTS=3`,任何语料都够不到)直接报 `citation_threshold_above_retrieval_limit`,不冒充材料缺口;preflight 阶段就做 manifest↔文件绑定,摘要或 `chunk_id` 对不上报 `corpus_entry_unbound`(在登录之前,不进半程失败);`source_path` 含 NUL 字节报 `corpus_entry_unusable`(`os.open` 抛 `ValueError`,已归一);未过**完整性门禁**的引用在**发起任何模型调用之前**就报 `citation_integrity_failed`;verdict 汇总阶段发现不支持的引用报 `citation_gate_failed`。这些都以至退出码 1 结束,不报「低分通过」。citation judge 的 CLAIM、QUOTE、SUBMISSION_FACTS 作为同一 JSON 对象里的独立数据值序列化,契约明确忽略其中的评分指令及伪造字段标签;该输入边界不等于已证明模型免疫注入。输出行标注 `judge=model` 与 `human_review=not_performed`,以便与将来的人工复核记录区分。verdict 默认写到**状态目录**(`$XDG_STATE_HOME` 为绝对路径时用它,否则 `~/.local/state`)下的 `ulticode/citation-verdicts-<随机>.json`:每次运行独立,且不落在 checkout 里,可用 `ULTICODE_CITATION_VERDICTS` 改路径;artifact claim 逐段 no-follow 打开并保留父目录 fd,锁、占用检查、临时文件、发布、回读及清理均相对该目录 fd 执行;模型调用期间或发布时父目录被替换不会重定向输出,显示路径失配则写入失败。发布以原子、不覆盖的硬链接完成,付费调用期间晚出现的目标也会被保留并报写入失败;失败清理仅对本次成功发布的目标核验 inode 后删除(核验与删除不是原子操作,不保证抵抗主动替换已发布文件的 writer)。无论哪种,目标位置都会在**付费调用之前**被独占占位,不可写或已被占用(包括 dangling symlink)即 `verdict_destination_unusable`;写盘同时生成 `<path>.meta.json` 附属文件,记录判定者(`judge=model`)、模型标签、语料、阈值、提交事实摘要,以及 `validated_corpus`——manifest 的 SHA-256 加上每条被判定文档的 id、version、已验证文件名、位置与内容摘要,使 verdict 与判定时的确切材料绑定、脱离本次会话仍可解释。密钥只从环境读取,不得写入日志或仓库;上面两条真实模型入口示例中的 `DEEPSEEK_API_KEY=...` 只是占位符,实际运行同样应先在环境中导出。

关键词 vs 向量的最小对照是**评测专用**的,不切换主路径,且需要一次性单机 Qdrant 与 `eval` 依赖组:

Expand Down
109 changes: 105 additions & 4 deletions services/agent/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -48,6 +48,40 @@ owner-reported total, so an account with a long submission history issues one re
request per page. The smoke scripts emit only fixed status labels and item
counts. They do not print response bodies, cookie names or values, tokens, submission source,
usernames, roles, tool names, source text, or model answer content.

### Answer-level evaluation (development split only)

`src/keyword_evaluation.py` measures retrieval only, so it records `citation_support` and
`answer_completion` as `DEFERRED` and `observed_behavior` as `not_measured` — a traceable
source id says where a fragment came from, not that it supports a conclusion. The answer-level
columns come from `src/answer_evaluation.py`, which adds an answer pass and a judging pass on
top of the same pinned keyword retrieval. The answer pass sees only the question and retrieved
fragments; expected/allowed/forbidden outcomes are withheld until judging. It returns explicit
retrieved chunk IDs as citations, and the judge/artifact use only those IDs, not every retrieval hit.

```bash
ULTICODE_ANSWER_EVAL=1 DEEPSEEK_MODEL=<model> uv run python e2e_answer_evaluation.py
Comment thread
DavidHLP marked this conversation as resolved.
```

Scope is enforced by the evaluator, not by the caller: it raises on any case outside
`development`, so `holdout` and `holdout2` stay sealed. The answer response must include `text`
and a `citations` array containing unique IDs from that case's retrieved fragments; empty citations
are valid when the answer cites nothing. An empty citation list records `citation_support` as
`not_applicable` rather than as a failed check. The run writes a run-scoped artifact under the
state directory carrying `scope=development_only`, `sealed_splits`, `judge=model`,
`human_review=not_performed` and every row — this is machine evidence, not human review.

Budget: two logical passes per case (answer, then judge), and the entry refuses
`DEEPSEEK_MAX_CALLS` below that plan before the first billed call. A stalled request — a
transport error or a timeout — is retried per call rather than restarting the batch, and a
retried attempt is billed, so each row records the attempts actually made in `model_calls`
rather than a fixed two. A protocol failure is not retried because the same input yields the
same shape. The destination artifact is reserved before the first billed call, an existing
artifact is never overwritten, it is published by rename, and the corpus and case file are
snapshotted once before the calls so the artifact identifies the material actually judged.
Tune `DEEPSEEK_TIMEOUT` (seconds, default 120) for a reasoning model that can exceed the
adapter's 30s default on one response.

## U02 boundary

U02 is preparing authorized-corpus retrieval and sourced analysis. The checked-in corpus is an
Expand All @@ -62,6 +96,71 @@ the deterministic sample slice only. The executable keyword evaluation is versio
module; authorized-corpus, vector-retrieval, real-model evaluation, and isolation evidence are
tracked in the U02 Linear tasks.

### Corpus override for the acceptance entry point

`e2e_citation_support_model.py` can analyse a corpus outside the repository instead of the
pinned sample, without editing either:

```bash
# `DEEPSEEK_API_KEY` must already be in the environment (exported by the operator or
# injected by the secret store) — it is never part of this command line, and without
# it the run stops at `deepseek_api_key_required`.
ULTICODE_CITATION_SUPPORT=1 DEEPSEEK_MODEL=<model> \
ULTICODE_CITATION_CORPUS_DIR=/absolute/path/to/corpus \
ULTICODE_CITATION_CORPUS_MANIFEST=/absolute/path/to/manifest.json \
uv run python e2e_citation_support_model.py
```

Contract — every violation is a fixed `FAIL reason=...` evidence line and exit 1:

- **Both variables or neither.** One set alone → `corpus_source_incomplete`. The run never
falls back to the pinned corpus: reporting evidence from a different corpus than the one it
was asked for is worse than stopping.
- **The root is a real directory**, not a symlink (`corpus_root_unusable`), and the manifest
declares nothing at all → `corpus_empty`, distinguished from a manifest that will not parse
(`corpus_manifest_unusable`).
- **Declarations come from the manifest.** Missing, unreadable, invalid or non-conforming
manifest → `corpus_manifest_unusable`. Every entry must declare exactly the material class
the run pins in source — `ACCEPTED_PERMISSION`, `ACCEPTED_SCOPE`, `ACCEPTED_SAMPLE_KIND` and
`ACCEPTED_ACCESS_SCOPE` (all four change together for authorised material) →
`corpus_declaration_unsupported`. The manifest is read **once**: its entries, the documents
they describe and the pinned class travel as one immutable snapshot through retrieval, the
worksheet and the verdict metadata, so replacing the file mid-run cannot leave them
describing different material.
- **One file per entry**, named by the manifest's own `source_path`, which must be a plain
name directly under the corpus directory — an absolute or external path is not the file that
was opened and is refused as `corpus_entry_path_not_relative`. The root is opened once with
`O_DIRECTORY|O_NOFOLLOW` and every entry relative to that descriptor — so neither the root nor
an entry can be swapped for a link between the check and the read: symlinked entry →
`corpus_entry_escapes_root`;
missing file →
`corpus_entry_missing`; two entries resolving to the *same file* — including two hard-link
names for one inode — → `corpus_entry_duplicate_source`, and copies under separate names →
`corpus_entry_duplicate_content`, compared on the canonical text so CRLF/LF and insignificant
whitespace variants are the same fragment (one fragment must never count as several
citations); unreadable, not UTF-8, empty, over `MAX_SOURCE_CHARS`, containing a NUL
byte in the declared path, or a read that stops short of EOF — a prefix that passes the
size check while a suffix stays unread → `corpus_entry_unusable`.
- **Positions are derived from the raw file**, spanning the first to the last non-blank
physical line, so leading blank lines are covered and content starting on line 3 reports
`lines 3-5` rather than `lines 1-5`. A manifest position that does not match →
`corpus_entry_position_mismatch`, so a citation cannot cite a location that does not exist.
- **The manifest is bound to its files at preflight**: entries whose `content_digest` or
derived `chunk_id` disagree with the document they describe are refused before the run
logs in → `corpus_entry_unbound`.
- **`ULTICODE_CITATION_REQUIRED_ROWS` above the retrieval limit is refused** rather than
reported as a material gap: retrieval caps at `MAX_RESULTS` (3), so a higher bar is
unreachable for any corpus → `citation_threshold_above_retrieval_limit`.
- With neither variable set, behaviour is exactly the pinned, manifest-gated sample corpus.

The policy is four pinned constants in `e2e_citation_support_model.py`. Authorised material
(DAV-58) changes all four together with its manifest in a reviewed commit; a corpus file cannot
grant itself a policy. The evidence line and verdict metadata report that pinned permission as
`corpus=...`, so a non-synthetic corpus is never labelled synthetic. The verdict sidecar also
carries `validated_corpus`: the manifest SHA-256 plus each judged document's id, version, verified
source name, position and content digest, so the verdicts are bound to the exact material rather
than to whatever the files hold after the run.

`data/keyword_cases.json` annotates every case with `required_evidence`, `answerable`,
`expected_behavior` (`cite`, `no_evidence`, or `refuse`), `allowed_behavior`, and
`forbidden_behavior`; the loader rejects a case missing any of them. `refuse` marks questions the
Expand All @@ -73,10 +172,12 @@ retrieval outcome, expected versus observed behavior, fabrication risk, tool cal
time. Any hit on a `refuse` case is recorded as a fabrication risk, and a case with no hit is not
counted as traceable because it has no citation to trace.

Retrieval facts and answer judgements are kept apart on purpose. `citation_support` and
`answer_completion` are recorded as `DEFERRED`: a traceable source id proves where a fragment came
from, not that it supports a conclusion, and deciding that needs the model or human pass tracked in
DAV-58.
Retrieval facts and answer judgements are kept apart on purpose. In the retrieval slice
`citation_support` and `answer_completion` are recorded as `DEFERRED`: a traceable source id
proves where a fragment came from, not that it supports a conclusion. The development split
fills those columns through `src/answer_evaluation.py`, which judges a generated answer rather
than reading them off document availability; the human support check tracked in DAV-58 remains
separate, and an artifact carrying `human_review=not_performed` is machine evidence only.

`src/retrieval.py` provides bounded keyword retrieval and source metadata. `src/sourced_analysis.py`
separates observed submission facts from hypotheses and only cites retrieved fragments. Java services
Expand Down
Loading
Loading