Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
23 changes: 23 additions & 0 deletions README_RU.md
Original file line number Diff line number Diff line change
Expand Up @@ -117,6 +117,29 @@ ownership; `[Codex]` готовит implementation work, а Codex APP меняе
| `[Codex]` | Implementation framing, code review, tests и release handoff. | Scoped execution package для repository work. |
| `[Thinkers OS]` | Thinker corpus, provenance, synthesis и pattern status. | Source-aware synthesis без выдуманной attribution. |

### Что можно показать снаружи

AI-OS — не просто библиотека prompts. Его ценность можно показать через
связанные рабочие поверхности:

| Поверхность | Что она даёт | Наблюдаемый масштаб |
|---|---|---|
| Семь ChatGPT Projects | Разделяют входящий поток, AI-подходы, решения, аналитику, LLM, разработку и корпус источников. | Явные владельцы, границы и handoff. |
| Analytics Factory | Ведёт от вопроса и data contract до расчёта, memo и QA. | Реестр из **22** аналитических методов. |
| Проверка поведения | Не смешивает consistency репозитория с поведением живого ChatGPT Project. | 99 детерминированных проверок, включая **22** регрессионных кейса; live-каталог — 45 кейсов × 3 запуска. |
| StreamDeck | Делает ежедневные маршруты и безопасные prompts доступными с двух устройств. | 16 переносимых профилей и 140 × 3 model-QA входов. |

22 метода — это не «22 примера ради количества», а словарь способов проверить
вопрос: изменение и структура, data quality и control, проверка объяснений и
взгляд вперёд. Полный реестр с предпосылками, ограничениями и владельцами
проверки находится в
[ANALYTICAL_TECHNIQUES.md](<ChatGPT/[Analytics]/Knowledge/ANALYTICAL_TECHNIQUES.md>).

В текущем репозитории нет вело-кейса: ни данных, ни готового анализа, ни
пользовательского сценария про велосипеды. Поэтому он не должен выглядеть как
уже готовая демонстрация. Такой кейс можно добавить отдельно в `[Analytics]`:
исходные данные → выбранные методы → проверяемые выводы → memo/визуализация.

Authoritative paths, instruction limits и AES applicability находятся в
[project registry](PROJECT_REGISTRY.md). Project packages разделены намеренно:
strategy discussion не должна молча становиться analytics calculation или
Expand Down
73 changes: 73 additions & 0 deletions benchmarks/live_behavioral/BASELINE_CONFIGURATION.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,73 @@
{
"observed_at": "2026-07-31",
"source_commit": "8d15f5a0f6c06a14bb4098d8d6dcbe7632457dcf",
"benchmark_version": "1.0.1",
"thinking_effort": "Medium",
"exact_model_identifier": "UNAVAILABLE",
"project_memory": "Default; not configurable in the observed UI",
"instructions_status": "PASS_7_OF_7_NORMALIZED",
"instructions_normalization": "Remove at most one trailing newline from both the UI and repository values.",
"instructions": {
"[AI OS]": {"ui_sha256": "9e262da88063e6569ad2084dcdf1fca11cd5d075d4b0d36841a581fdea2ab926", "repo_sha256": "4dcf4b2dbb7d14f1f52f7778dce87c0ad4310666624282058d4c5925ed60bbd8", "normalized_sha256": "9e262da88063e6569ad2084dcdf1fca11cd5d075d4b0d36841a581fdea2ab926", "match": true},
"[Thinking]": {"ui_sha256": "491c601da609fb2f0dc770d96f52d89aa8d5a536c5729e2eef50f661673b63fe", "repo_sha256": "5dde30d7c0364e3eec10a38e0c3a34a39e768d9a1968d493c265da3d38549dab", "normalized_sha256": "491c601da609fb2f0dc770d96f52d89aa8d5a536c5729e2eef50f661673b63fe", "match": true},
"[Analytics]": {"ui_sha256": "cd977b8eaeaeb8ed93ef4a5365fa8a2017177b9177a21fc18518a3f1405f74b7", "repo_sha256": "b51574737afb37fb001d860908aae78f3856f0c2e6915e7509c71984852be7fd", "normalized_sha256": "cd977b8eaeaeb8ed93ef4a5365fa8a2017177b9177a21fc18518a3f1405f74b7", "match": true},
"[LLM]": {"ui_sha256": "443c323f2ac751f0dc17cc9ce4d59bdf389444e96a49771dca56b63c16b17307", "repo_sha256": "e7165b25650406f64b61a2ac40de2b871e3157dceb1ead37fe86ed1196552db2", "normalized_sha256": "443c323f2ac751f0dc17cc9ce4d59bdf389444e96a49771dca56b63c16b17307", "match": true},
"[Codex]": {"ui_sha256": "17789e9da1ba27a91ae3536e651374541f54f3db46678e198cb50bedb79ec86e", "repo_sha256": "793d49ef0f9ac74770efe221f9205d7af51bf3a2aa22829b23a673acf5dd6593", "normalized_sha256": "17789e9da1ba27a91ae3536e651374541f54f3db46678e198cb50bedb79ec86e", "match": true},
"[Inbox Router]": {"ui_sha256": "d60e8776fff384e6256aa07a715f539ecdd44baf646c41816c613766c543210b", "repo_sha256": "1ab0ebcc0681e50b28a775be824b66abe76fc56b6b9903ccbaad804a6580c6ca", "normalized_sha256": "d60e8776fff384e6256aa07a715f539ecdd44baf646c41816c613766c543210b", "match": true},
"[Thinkers OS]": {"ui_sha256": "3582a058bb74726266e8d4d4f7827e9d775b8b77ffb79fb642e34f02dd1aca7a", "repo_sha256": "5e419ca1be9b3c89ceb90e5d89d316a431f679d16bfbe92ea8729da876ab946d", "normalized_sha256": "3582a058bb74726266e8d4d4f7827e9d775b8b77ffb79fb642e34f02dd1aca7a", "match": true}
},
"knowledge_status": "FILENAMES_PASS_SERVER_BYTES_UNVERIFIED",
"knowledge_files": {
"[AI OS]": {
"AIOS_01_ROUTING_AND_WORKFLOW.md": "95ea21d47d645e961ae081bdd8c81467d865d85901bb77d478bee0db16ed2e6d",
"AIOS_02_GOVERNANCE_AND_EVIDENCE.md": "dbb3bae0c8b2e63aa73220e68f0c9bcc668ab79e7945c3af39f89b5a2144f975",
"AIOS_03_HANDOFF_AND_SMOKE_QA.md": "58eb37b5e3a7c17a6ece68f051b8ccd6cac85823623b0f39a29a6f80a8001b3c",
"AIOS_04_GOAL_PACKS_AND_COMMAND_SURFACE.md": "96884282ae082960f0a5fc40a5b5d1083a07c6e168b47ea9c4972b37e61b021c",
"AIOS_05_SUPERVISED_AGENT_LOOPS.md": "a0b48541de20ca000a341cdf50819908317447b6ab80ed21009c83748516fea6",
"AIOS_06_CROSS_PROJECT_AI_EVALS.md": "43834bac395123909bbff82f9108df079ad0b7413cee895c50c4dae674d616b9"
},
"[Thinking]": {
"THINKING_01_WORKFLOW_AND_DECISIONS.md": "53e78a4b08701b0f838033c798218806d20eb4334bcb8d42fb8f5373191049c7",
"THINKING_02_JUDGE_REVISOR_RISK.md": "3bd89fd3437b5a7b346e0ae2569fea41c9dc6e4059e64a1371628b98d13e438f",
"THINKING_03_ROUTING_AND_TEMPLATES.md": "a44abc11da9e88eae95fda8fb543b7fbfe2e18d4c6d157003020c0adae422315",
"THINKING_04_THINKERS_SYNTHESIS.md": "834476e2b2ecda044540b80ec5a5fe2a49b3e43d8f94c6cffb84c4595ffa3e36"
},
"[Analytics]": {
"ANALYTICS_01_CORE_WORKFLOW.md": "4a1ca3fe212d595a700ccfaee58e1e3486d18b4fca52fd57860b43f7fb438ee6",
"ANALYTICS_02_DATA_CONTRACTS_AND_MARTS.md": "b11d4db7cdd4c72cec7564dc9adcf81787f2b8b1bce81f8957bb3b68b7efd7ee",
"ANALYTICS_03_TECHNIQUES_AND_CHARTS.md": "7ca4449e5fcb9822c3c945289f4e2b7834c5d86c275180e28cb481457df05af0",
"ANALYTICS_04_MEMO_AND_TEXT_STANDARDS.md": "3f305574d51fdac5c6d42e131235f7adba3476f0d99219c6e2b6746e72bb805d",
"ANALYTICS_05_QA_GOVERNANCE_ROUTING.md": "b3a5f78e13454e72cc5cba07b297fea035d12f484bbf8069f2b762f6475ab914",
"ANALYTICS_06_TEMPLATES.md": "d2815aa82f1164c100aca4706ee3f8eab707fd64f77e61b3feb622d953409dff",
"ANALYTICS_07_CODEX_HANDOFF_OPTIONAL.md": "8f23c9690e92f72a1a8a6f4fc097541603118026de99a0059da0d52b4da7397e"
},
"[LLM]": {
"LLM_01_ROUTING_AND_MODEL_SELECTION.md": "92c8fa2dcdc12686b92b01308c183a34f2b2e5e8bfb017810c99c23f443a0b49",
"LLM_02_PROMPT_LIBRARY_AND_REGISTRY.md": "d55e95ab252c0eae2671f5c58161937b5a4c979c68abe6c1a49544016406b669",
"LLM_03_QUALITY_GATES_AND_EVAL.md": "030120ac3a62e0b9abb270cf4c91b0e4a4dc4a1dfde06d8a28adbd0c4cf17201",
"LLM_04_WORKFLOWS_AND_HANDOFF.md": "0df3d23c84ad25513da027e7a89c2ed6e3d0856cef01867e011f53aae2481c67",
"LLM_05_CONTEXT_ENGINEERING.md": "901fd59eab0945bc4d02ba160f7f568001b816e945bc32093da4deb84bed0d87",
"LLM_06_LOCAL_AI_EXPERIMENTS.md": "0c3365042946e299e3284dc72789496327f106b4541051af841b192657b041e1"
},
"[Codex]": {
"CODEX_01_TASKS_AND_HANDOFF.md": "fb04e5bed524b6b12d534a0d59302d3bdf0962cd2620393ac814eaa1ec2c05b8",
"CODEX_02_EXECUTION_AUTONOMY_REPORTING.md": "15ca9f00bc54209932c5e87bbbd9106f8e7c27795b6c21306c12e75d8f656f39",
"CODEX_03_TESTING_ACCEPTANCE_RELEASE.md": "8665957970d11a205c42001f27721af056f987800a05fd20032eec2527da26d3",
"CODEX_04_IMPLEMENTATION_WORKFLOWS.md": "bd382efb19903893f4bbed2ad473a8f754a4e7a135fee18b8540e108d39c31b9",
"CODEX_05_AGENT_REFERENCES.md": "65bd1d7b94cf7eee852531574a70e239840f2f288f6e00d677ad03b196894fe7",
"CODEX_06_AI_CODING_DISCIPLINE.md": "280c295e9a4ed84f36cc313d6055b93e2f0802b83061c2103215f3874bb2bbfe"
},
"[Inbox Router]": {
"INBOX_01_ROUTING_WORKFLOW.md": "b0240b47893a7b69232ffbde12e7613e0d599fb791ad1c2f024674aa11c90290",
"INBOX_02_HANDOFF_QA_ANTI_PATTERNS.md": "c52028a14f310ee8ef5631303498a9a03657b484653a63a3e3b62440f49e8050"
},
"[Thinkers OS]": {
"THINKERS_OS_01_PORTFOLIO_AND_CORPUS.md": "2727b5a9236a241b498b2011999312f20a19ccc74dcadcf71efce60063fad13b",
"THINKERS_OS_02_ARTIFACTS_AND_SYNTHESIS.md": "92b2f3ec21a39f5c43426ab8e935b92602cbe96d89c314bf3b37993ebce2f339"
}
},
"knowledge_ui_evidence": "All 33 declared bundle filenames were observed in their target Project. AI OS also retains the governed KB__ files documented as intentional external Knowledge.",
"knowledge_server_byte_equivalence": "UNVERIFIED",
"duplicate_policy": "Existing matching filenames were preserved. Duplicate upload prompts were skipped; no Project source was deleted and no duplicate was added.",
"configuration_comparability": "PASS_WITH_DOCUMENTED_EXTERNAL_RUNTIME_LIMITATIONS"
}
7 changes: 7 additions & 0 deletions benchmarks/live_behavioral/BENCHMARK_CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
# Benchmark Changelog

## 1.0.1

Before baseline, live UI inspection showed that `[Thinking]` preserves the repository's final newline while other Project textareas may strip it. Version 1.0.0 incorrectly normalized only the repository side and therefore misclassified `[Thinking]` as stale.

Version 1.0.1 removes at most one trailing newline from both values before exact comparison. No cases, prompts, rubric criteria, weights, floors, hard-fail rules, thresholds, holdout content or sealed holdout hash changed. Version 1.0.0 was never baselined and is not evidence.
19 changes: 19 additions & 0 deletions benchmarks/live_behavioral/CAPABILITY_MATRIX.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
# Live Benchmark Capability Matrix

Assessment date: 2026-07-31. Browser evidence was observed in the authenticated ChatGPT Project UI.

| Capability | Status | Evidence | Limitation |
|---|---|---|---|
| Actual ChatGPT Project access | supported | Authenticated Project home/chat pages opened for all seven target Projects. | Session authentication is user-owned. |
| Project selection | supported | Stable Project URLs and visible Project names were observed. | Project IDs are external runtime identifiers. |
| Exact Instructions loading | supported | Settings textarea can be read and updated; hashes can be compared with repository Instructions after removing at most one trailing newline on both sides. | `[Inbox Router]` was stale and `[Thinkers OS]` was empty at discovery; they must be synchronized before the valid baseline. |
| Exact Knowledge loading | supported | Required repository files can be uploaded through the Project Sources UI and filenames can be enumerated afterward. | The UI does not expose post-ingestion bytes, so server-side byte equivalence remains UNVERIFIED; source file hashes and observed filenames provide provenance. |
| Raw response capture | supported | Full visible assistant response, prompt, chat URL and hashes can be captured from each saved Project chat. | Raw captures remain local; repository copies must be anonymized. |
| Repeated identical prompts | supported | Fresh Project chats can be created repeatedly with the same prompt. | Product-side sampling is not controllable. |
| Baseline/candidate separation | supported | Configuration hashes and fresh chat URLs distinguish phases. | Same external account and product runtime are used. |
| Exact model pinning | unsupported | Project UI exposes `Medium` thinking effort but no exact model identifier in the composer. | Model comparability is UNVERIFIED if the product changes routing. |
| Context preservation | supported | Every run starts from the same Project home in a new empty chat. | Project memory is `Default` and cannot be changed in the UI. |
| Holdout isolation | supported | Holdout prompts are generated and sealed before baseline, then opened only by Runner/Evaluator after candidate selection. | Procedural isolation on one physical system; independent enforcement is UNVERIFIED. |
| Independent evaluator | UNVERIFIED | Runner, evaluator and final judge use separate artifacts and hashes. | One physical system performs all roles; residual risk is self-evaluation bias. |

No Level C improvement claim is allowed until actual baseline and candidate Project runs are captured and compared.
15 changes: 15 additions & 0 deletions benchmarks/live_behavioral/COVERAGE.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
# Coverage Matrix

Public development benchmark: 45 cases × 3 fresh-chat runs = 135 live Project runs per phase. A separately sealed holdout is not included in public counts.

| Project / route | Positive | Negative | Cross-project | Readability simple | Readability complex | Adversarial | Public cases | Runs / phase |
|---|---:|---:|---:|---:|---:|---:|---:|---:|
| [Inbox Router] | 2 | 1 | 1 | 1 | 0 | 2 | 7 | 21 |
| [AI OS] | 2 | 1 | 1 | 1 | 1 | 1 | 7 | 21 |
| [Thinking] | 2 | 1 | 1 | 1 | 1 | 2 | 8 | 24 |
| [Analytics] | 2 | 1 | 0 | 0 | 1 | 1 | 5 | 15 |
| [LLM] | 2 | 1 | 1 | 1 | 1 | 1 | 7 | 21 |
| [Codex] | 2 | 1 | 1 | 1 | 0 | 1 | 6 | 18 |
| [Thinkers OS] | 2 | 1 | 0 | 0 | 1 | 1 | 5 | 15 |

Set totals: routing 21; response quality 5; readability 10 (5 simple, 5 material complex); adversarial 9. Every tested Project has three core routing cases: two positive and one negative. The nine critical hard-fail classes are each represented by at least one tagged public case.
24 changes: 24 additions & 0 deletions benchmarks/live_behavioral/DISCOVERY_EVIDENCE.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,24 @@
{
"observed_at": "2026-07-31",
"source_commit": "8d15f5a0f6c06a14bb4098d8d6dcbe7632457dcf",
"instructions": {
"[AI OS]": {"ui_sha256": "9e262da88063e6569ad2084dcdf1fca11cd5d075d4b0d36841a581fdea2ab926", "repo_ui_normalized_sha256": "9e262da88063e6569ad2084dcdf1fca11cd5d075d4b0d36841a581fdea2ab926", "match": true},
"[Thinking]": {"ui_sha256": "5dde30d7c0364e3eec10a38e0c3a34a39e768d9a1968d493c265da3d38549dab", "ui_normalized_sha256": "491c601da609fb2f0dc770d96f52d89aa8d5a536c5729e2eef50f661673b63fe", "repo_ui_normalized_sha256": "491c601da609fb2f0dc770d96f52d89aa8d5a536c5729e2eef50f661673b63fe", "match": true},
"[Analytics]": {"ui_sha256": "cd977b8eaeaeb8ed93ef4a5365fa8a2017177b9177a21fc18518a3f1405f74b7", "repo_ui_normalized_sha256": "cd977b8eaeaeb8ed93ef4a5365fa8a2017177b9177a21fc18518a3f1405f74b7", "match": true},
"[LLM]": {"ui_sha256": "443c323f2ac751f0dc17cc9ce4d59bdf389444e96a49771dca56b63c16b17307", "repo_ui_normalized_sha256": "443c323f2ac751f0dc17cc9ce4d59bdf389444e96a49771dca56b63c16b17307", "match": true},
"[Codex]": {"ui_sha256": "17789e9da1ba27a91ae3536e651374541f54f3db46678e198cb50bedb79ec86e", "repo_ui_normalized_sha256": "17789e9da1ba27a91ae3536e651374541f54f3db46678e198cb50bedb79ec86e", "match": true},
"[Inbox Router]": {"ui_sha256": "922772b4c00020cb2b78b9bf89822541cb9884338e17a1778fc13af4e0a1fbe1", "repo_ui_normalized_sha256": "d60e8776fff384e6256aa07a715f539ecdd44baf646c41816c613766c543210b", "match": false},
"[Thinkers OS]": {"ui_sha256": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855", "repo_ui_normalized_sha256": "3582a058bb74726266e8d4d4f7827e9d775b8b77ffb79fb642e34f02dd1aca7a", "match": false}
},
"knowledge_filenames": {
"[AI OS]": ["AIOS_01_ROUTING_AND_WORKFLOW.md", "AIOS_02_GOVERNANCE_AND_EVIDENCE.md", "AIOS_03_HANDOFF_AND_SMOKE_QA.md", "AIOS_04_GOAL_PACKS_AND_COMMAND_SURFACE.md", "AIOS_05_SUPERVISED_AGENT_LOOPS.md", "AIOS_06_CROSS_PROJECT_AI_EVALS.md", "KB__00_INDEX.md", "KB__01_NAVIGATION.md", "KB__02_CONTENT.md", "KB__03_WORKFLOWS_TRACEABILITY.md", "KB__04_SMOKE_QA.md", "KB__05_CANONICAL_CONCEPTS.md", "KB__06_OPERATIONAL_FRAMEWORKS.md", "KB__07_PATTERNS_AND_FAILURES.md", "KB__08_USE_CASES_FOR_SERGEY.md", "KB__CARD_SCHEMA.md", "KB__CHANGELOG.md", "KB__CONFIDENCE_RULES.md", "KB__DEDUPLICATION.md", "KB__PROMOTION_GATES.md", "KB__RELEASE_MANIFEST.md", "KB__RETRIEVAL_QA.md", "KB__REVIEW_QUEUE.md", "KB__USE_CASE_ROUTING.md"],
"[Thinking]": ["THINKING_01_WORKFLOW_AND_DECISIONS.md", "THINKING_02_JUDGE_REVISOR_RISK.md", "THINKING_03_ROUTING_AND_TEMPLATES.md", "THINKING_04_THINKERS_SYNTHESIS.md"],
"[Analytics]": ["ANALYTICS_01_CORE_WORKFLOW.md", "ANALYTICS_02_DATA_CONTRACTS_AND_MARTS.md", "ANALYTICS_03_TECHNIQUES_AND_CHARTS.md", "ANALYTICS_04_MEMO_AND_TEXT_STANDARDS.md", "ANALYTICS_05_QA_GOVERNANCE_ROUTING.md", "ANALYTICS_06_TEMPLATES.md", "ANALYTICS_07_CODEX_HANDOFF_OPTIONAL.md"],
"[LLM]": ["LLM_01_ROUTING_AND_MODEL_SELECTION.md", "LLM_02_PROMPT_LIBRARY_AND_REGISTRY.md", "LLM_03_QUALITY_GATES_AND_EVAL.md", "LLM_04_WORKFLOWS_AND_HANDOFF.md", "LLM_05_CONTEXT_ENGINEERING.md", "LLM_06_LOCAL_AI_EXPERIMENTS.md"],
"[Codex]": ["CODEX_01_TASKS_AND_HANDOFF.md", "CODEX_02_EXECUTION_AUTONOMY_REPORTING.md", "CODEX_03_TESTING_ACCEPTANCE_RELEASE.md", "CODEX_04_IMPLEMENTATION_WORKFLOWS.md", "CODEX_05_AGENT_REFERENCES.md", "CODEX_06_AI_CODING_DISCIPLINE.md"],
"[Inbox Router]": ["INBOX_01_ROUTING_WORKFLOW.md", "INBOX_02_HANDOFF_QA_ANTI_PATTERNS.md"],
"[Thinkers OS]": ["THINKERS_OS_01_PORTFOLIO_AND_CORPUS.md", "THINKERS_OS_02_ARTIFACTS_AND_SYNTHESIS.md"]
},
"knowledge_byte_equivalence": "UNVERIFIED",
"knowledge_limitation": "The Project UI exposes filenames and accepts exact local uploads but does not expose post-ingestion file bytes."
}
11 changes: 11 additions & 0 deletions benchmarks/live_behavioral/HOLDOUT_MANIFEST.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,11 @@
{
"status": "sealed_before_baseline",
"case_count": 7,
"sealed_sha256": "e085abd624c1984d8c9d8bc0ac4812784ab0e61fa3f93bf25547668f1c034021",
"algorithm": "aes-256-gcm",
"location": "untracked local runtime artifact",
"optimizer_disclosure": "NOT DISCLOSED",
"open_stage": "after candidate selection",
"independent_enforcement": "UNVERIFIED",
"limitation": "Procedural isolation on one physical system; residual risk is self-evaluation bias."
}
19 changes: 19 additions & 0 deletions benchmarks/live_behavioral/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
# SUPERManager Live Behavioral Benchmark

This benchmark evaluates actual responses from the seven AI-OS ChatGPT Projects. Static repository checks do not substitute for live runs.

Workflow:

1. freeze benchmark, rubric, public cases, coverage and sealed holdout hashes;
2. synchronize the source-baseline Instructions and Knowledge into the actual Projects;
3. run every public case three times in a fresh Project chat;
4. capture prompt, full raw response, Project URL, configuration hash, model condition and response hash;
5. evaluate raw responses without Optimizer rationale;
6. make at most five bounded configuration iterations;
7. repeat the same full run procedure for the selected candidate;
8. open the sealed holdout only after candidate selection;
9. apply the frozen final gate and create a separate PR without merging.

Raw responses and the sealed holdout stay outside the repository. The PR contains anonymized cases, response samples, aggregate results and hashes without private data.

Independent evaluation: UNVERIFIED. Residual risk: self-evaluation bias.
Loading
Loading