Verify it yourself
Proof — verify our claims yourself
Section titled “Proof — verify our claims yourself”We sell falsifiability, so the selling is machine-checked. Every public claim binds to resolvable evidence; a skeptic can reproduce every check below on a fresh checkout. This page is itself generated from those sources and fails CI if it drifts.
What this prevents
Section titled “What this prevents”Each row is a failure mode, not a feature — and each cites the ledger claim that backs it. A row whose claim stopped resolving would fail this page’s own generator, so the table cannot outlive its evidence.
| Failure this prevents | How | Backed by |
|---|---|---|
| A rule says MUST, and nothing anywhere enforces it — so the rule reads as a guarantee while being honour-system. | Each rule declares enforced_by:, and the check resolves it: a validator reachable from no taskfile, workflow, or hook manifest counts as unwired, not as covered. Undeclared counts as uncovered. |
enforcement-coverage-resolved ⟳ |
| A shipped artefact carries a hidden-Unicode or instruction-smuggling payload, and reaches consumers because only the source tree was scanned. | Source and the condensed projection are scanned in CI; a finding blocks the release before publish, not merely the merge. | shipped-artifacts-hidden-instruction-scanned ⟳ |
| A count in public prose drifts from the source, and the marketing number quietly stops being true. | Counts are generated from source and drift-checked; the build fails on a count-shaped prose mention that disagrees, or on two different numbers for the same artefact kind. | skill-count ⟳ |
| A host-coverage claim understates or overstates what actually works — for months, on the one surface everybody reads. | The number is pinned by a test over real detection, and the test is re-run as the claim evidence. | host-agent-count ⟳ |
| A non-coding domain skill implies proven correctness it never had, because it was forged on a different domain and never checked against one. | Domain skills are labelled unvalidated until they pass a sourced domain-truth fixture, and the validated count is ratcheted so it cannot quietly fall. | domain-soundness-scoped ⟳ |
| Behavioural-eval coverage regresses as skills are added, and the suite looks healthy because only the absolute number is reported. | Coverage is measured per tier and ratcheted in CI, so it can only rise; the gap is published rather than implied away. | eval-coverage-ratcheted ⟳ |
⟳ — the claim carries exec: evidence: CI re-runs the command and
compares its exit code, so a stale row turns the build red rather than
ageing quietly into marketing.
See it run (< 60s, real output)
Section titled “See it run (< 60s, real output)”
The recording is of the exact commands in § 5, nothing staged. Its
commands are re-executed in CI (.github/workflows/proof-demo.yml), so a
demo that showed something broken would turn CI red — the recording
cannot drift from current behavior.
1. Every public claim binds to evidence
Section titled “1. Every public claim binds to evidence”Rendered from docs/CLAIMS.md. A <!-- claim:ID --> marker in
README/docs must resolve to a backed entry here with a resolving
evidence pointer, or task check-claims fails the build.
| Claim | Kind | Evidence | Resolves |
|---|---|---|---|
| The two surfaces the provider-lifecycle contract obliges to agree on an adapter tier — the adapter header and the xml example — do agree, and the surface that read stale is no longer hand-written: § 5 is generated, so the drift this claim was opened over cannot recur. | qual | docs/contracts/provider-lifecycle.md#Current tier assignment |
✅ |
adversarial-review structures its critique as branching exploration with explicit pruning rather than a single pass, following the Tree-of-Thoughts formulation. |
qual | https://arxiv.org/abs/2305.10601 (2026-08-11) |
✅ |
The .augment-plugin/ manifest version is the package version, not an independent plugin-API version, and every version-bearing file the release workflow triggers on is read by a job in that workflow. |
qual | src/scripts/lint_marketplace.ts#check_augment_manifests |
✅ |
analysis-autonomous-mode runs iterative self-critique between steps rather than only at the end, following the Self-Refine formulation. |
qual | https://arxiv.org/abs/2303.17651 (2026-08-11) |
✅ |
bug-analyzer verifies each candidate root cause against a concrete trigger before reporting it, following the Chain-of-Verification formulation — which is the mechanism behind its “never invent issues” constraint. |
qual | https://arxiv.org/abs/2309.11495 (2026-08-11) |
✅ |
| The release process is documented as an inheritable runbook + succession doc, and the project’s bus-factor (trailing-90-day distinct human reviewers) is tracked and reported truthfully — currently 1, not implied to be more. The doc separately reports the distinct MERGER count (2, one of them an unreviewed self-merge) so the reviewer figure cannot be inflated by conflating the two. | qual | docs/succession.md#trailing 90 days |
✅ |
On the pre-registered 2-arm retrieval benchmark (18 hand-verified code-structure questions across 3 real repos, ground truth hash-bound before the run, deterministic, zero model calls), the native code graph scored mean recall 0.365 vs grep 0.797 on the 15 graph-shaped questions (delta -43.2 pp against a pre-declared +10 pp win threshold) and 0.111 vs 0.833 on the negative controls. HONEST NULL — measured root cause: TS arrow-function exports produce no symbol nodes (170 TS vs 13,428 PHP symbol nodes on same-shaped repos) and string-keyed dynamic consumers have no static edge. Consequence bound: code_graph.enabled stays false BY DEFAULT. NOT permanently, and the removal schedule this line used to carry was withdrawn on 2026-08-15 (docs/MIGRATION.md, the code_graph row) — the payload it existed for had already shipped when the parser pair moved to devDependencies. SCOPE OF THESE FIGURES, corrected 2026-08-26: they were measured on 2026-07-28 against a build predating the extractor repair of 2026-08-22, so they describe a build that no longer exists. The claim is retained rather than dropped because the measurement was real and no re-measurement has replaced it; road-to-inbox-harvest-2026-08-f-code-graph-evidence-refresh 3.1 is the re-run and it is blocked on benchmark inputs this repository does not hold. Reopen on that re-run, or on external evidence (a consumer case the graph answers and disciplined grep cannot). RE-MEASUREMENT ATTEMPTED 2026-08-26 and deferred: the four SHA-256-pinned question files and all three registered corpus clones read absent on the machine that attempted it, so the run could not start. The obligation is transferred, not dropped - agents/roadmaps/stubs/road-to-code-graph-benchmark-rerun.md carries the criterion verbatim with a five-reading probe. Silence here would read as ‘nobody tried again’. RE-MEASUREMENT PERFORMED 2026-08-28 on a DIFFERENT, SMALLER, NON-COMPARABLE corpus, because the pinned inputs re-probed absent again on that date: internal/bench/reports/code-graph-vs-grep-inrepo-2026-08-28.md, pre-registered before the run at internal/bench/code-graph/PREREGISTRATION-inrepo-2026-08-28.md, measured at commit c454648af which postdates the repair. THAT RUN IS NOT A REPLACEMENT FOR THESE FIGURES AND NO DELTA MAY BE COMPUTED BETWEEN THE TWO - it used three in-repo TypeScript subtrees instead of the three registered private repositories, 16 questions sharing no item with these 18, and per-class bars instead of a single aggregate; three independent variables moved at once. It is recorded here as a POINTER so this entry is not the only answer a reader can find, and it deliberately carries no claims entry of its own: the kind enum is {quant, qual, comparative} and none of the three makes two recall figures incomparable by construction, so a second entry would put two subtractable numbers on this surface. On its own terms that run reports ZERO of four graph-shaped classes meeting the pre-registered win criterion, with one class (path-between) VOID and the literal-string negative-control floor failed by construction against a symbol index; it withholds an overall engine verdict for exactly that reason. TWO CORRECTIONS TO THAT POINTER, 2026-08-29. (a) THE PUBLISHED ROOT CAUSE OF THE VOID CLASS WAS FALSE. That report said ‘both arms returned the empty set on every question in this class’. Only the grep arm did - a word-boundary search for a token containing ’ -> ’ matches no text. The GRAPH ARM ANSWERED all three questions and the v1 scorer discarded the answer, comparing each returned symbol against the entire probe string ‘cmdBuild -> getParser’. Two further v1 scorer defects: it never invoked the shipped path <a> <b> verb, and it counted unresolved symbol: pseudo-nodes as files, which is the sole reason its callers class was ruled NULL with recall tied at 1.000/1.000. v1’s NUMBERS ARE NOT RETRO-EDITED - they were faithful to v1’s own registration - and v1’s report now carries the correction inline. (b) THE COMMIT POINTER c454648af IS UNREACHABLE FROM main: it exists only on a local branch and resolves to nothing in a fresh clone. The reachable equivalent is 6a5670b78 (the squash merge of PR #1705, on main), at which all three measured roots and the corpus file are byte-identical and the scoring functions are unchanged, so that run reproduces there. Measured trees, which survive a squash merge: src/scripts/code_graph e55fc87c7081359e1c0e2bf8263f17358de2e315, src/shared dcdc68073e0f2340b3734678cc4e82f34c241dba, src/scripts/ai_council 778a493ba915795159adcb0bb789e93b89d6a86a. RE-MEASUREMENT v2 PERFORMED 2026-08-29 after those defects were confirmed by direct execution: internal/bench/reports/code-graph-vs-grep-inrepo-v2-2026-08-29.md, pre-registered at internal/bench/code-graph/PREREGISTRATION-inrepo-v2-2026-08-29.md with bars identical to v1’s, 19 questions, the path verb used for the path class, symbol: pseudo-nodes no longer counted as files, and the literal probes moved to a separate capability-boundary class with no floor derived from them. IT IS NOT COMPARABLE TO THIS ENTRY AND NOT COMPARABLE TO v1 EITHER, and carries no claims entry of its own for the same reason v1 does not. On its own terms v2 reports ZERO winning classes: callers TIE, path-between TIE (+8.3 pp against a +10 pp bar, the graph exact at recall 1.000 precision 1.000 and the repaired grep arm at 0.917), transitive-impact NULL, references NULL. A path-between delta near +89 pp is obtainable and is an ARTIFACT - it repairs the graph arm and leaves grep on v1’s broken probe. ADR-246’s reopen trigger was evaluated explicitly against v2 and DOES NOT FIRE; no setting default moved and no dependency moved between devDependencies and dependencies. The figures in THIS entry still describe the pre-repair build and are still the only measurement of the registered corpora. |
quant | internal/bench/reports/code-graph-vs-grep.md#Verdict — NULL, decisively |
✅ |
| 205 commands. | quant | exec:check_artefact_count_messaging -> 0 |
✅ |
On the first post-fix behavior-conformance measurement (round 5, /analyze:conformance --limit 30, run 2026-08-07, every violation split by its own timestamp against the 2026-08-06 carrier merge), both BLOCKING pre_tool_use guards eliminated their classes — unauthorized irreversible git ops 8 → 0, evaluator prompt pre-loading its verdict 1 → 0 — while neither advisory carrier did: language-mirror violations fell 555 → 19 under advisory state injection at user_prompt_submit (−96.6%, not zero) and verification-claimed-on-empty-output fell 4 → 1 at post_tool_use. Advisory reduces massively; only blocking eliminates. Scope bounds, inseparable from the numbers: the post-fix corpus is ONE session (~600 assistant turns), so this is a recorded prior, not a law; the language pin was verified PRESENT on the violating post-fix turns (which is why an UNBOUNDED re-pin — restating it on every tool call — stays refused; what shipped 2026-08-20 instead is a BOUNDED re-emit, once per 150 tool calls, against a red baseline this round-5 corpus could not produce and whose own revisit clause called for: 11 English replies to a German user, all of them 179+ tool calls past the pin, none below. Corrected here because the unqualified earlier wording — “higher injection frequency is refused as a fix” — is a claim the shipped code contradicts; the reasoning, the measured distances, and the accepted false-fire cost are in agents/settings/contexts/reminder-injection-verdict.md); and the 555 pre-merge count is partly contaminated by the synthetic-turn mis-pin fixed in the same round (isSyntheticPrompt). NOT PRE-REGISTERED, and it could not have been: this is a post-hoc audit reading taken after the carrier landed, so no bar existed to freeze before the data. What stands in for pre-registration is that the one choice a post-hoc split could game — where before ends and after begins — was not chosen by the measurement: the boundary is the carrier merge’s own commit timestamp (2026-08-06) and every violation is assigned by its own timestamp against it. The detector is deterministic and re-runnable over the same transcript store, so the numbers are reproducible rather than attested. Read it as a recorded prior, never as a pre-registered result; the sibling claim that DOES carry a frozen bar is the scoped-rule absence experiment, pre-registered precisely because its data does not exist yet. |
quant | src/domains/analysis-workbench/analyze/conformance/command.md#Both blocking carriers reached zero |
✅ |
A MEASURED-BUT-NOT-SHIPPED experiment — the thin rule projection reduced eager rule load 78,513 → 13,881 GPT-tokens (whole always-loaded projection 98,529 → 33,897, ~65.6%), but FAILED the quality gate (thin win-rate 36.2% vs required 48%) and does not ship; it un-defers only behind discipline_profile: essential. Shipped behavior does NOT include this reduction. Method: agent-config benchmark over the pinned token baseline; the baseline is the honest “what the user pays if everything loads eagerly”, NOT a synthetic full-corpus strawman (council Q4); quality gate per the Phase-0 paired judge run. |
quant | internal/bench/reports/token-baseline.json#eager_rule_load |
✅ |
An eligible mid-flight cli failure with a constructible api twin loses zero council seats WHEN THE PROJECTED-SPEND GATE PERMITS THE RETRY — the seat answers over the api rung instead of dropping out of the pass. A retry the budget REFUSES is outside the claim: the original failure stands, the seat is absent, and fallback_skipped: cost_budget says so. |
qual | exec:vitest run tests/scripts/ai_council/council_cli.test.ts -> 0 |
✅ |
A council member’s unparseable answer is separable from a member that found nothing — 2/7 parse_failed, 1/7 empty, 4/7 parsed over the seven recorded answers in tests/fixtures/council-parse-corpus/, which is a fixture-corpus denominator and NOT live traffic; reproduce with ./scripts-run src/scripts/council_parse_rate. |
quant | exec:vitest run tests/scripts/ai_council/parse_corpus.test.ts -> 0 |
✅ |
The first cross-vendor parity pass (5 orchestration-corpus tasks × 2 vendors × 3 repeats, identical prompts via the council transport, $0.16) measured real per-host finding-count differences — claude-sonnet-4-5 surfaced ~2× the findings of gpt-4o on the multi-file analysis task (median 11 vs 5) while both vendors were identical on the planted hollow-implementation task (2 vs 2) and perfectly silent on the clean-code negative control (0 vs 0, no spurious findings). The per-task finding_floor values are calibrated from the cross-host lower envelope and the gate is armed. |
quant | internal/bench/reports/parity-count.json#min over hosts of median |
✅ |
The scoped-projection default for new installs ships 228 of 299 skills (untagged core plus engineering/maintainer packs), an approximately 25% reduction of the skill-catalog surface (a reduction of 71 projected entries; the token figures measured 2026-07-27 at the then-283-skill catalog were about 577k to about 428k approximated tokens and are NOT rescaled here). Both figures are generated, not typed: reproduce them with ./scripts-run src/scripts/count_scoped_projection, which partitions the canonical skill catalog with the same predicate install.ts applies when it prunes a real tree. |
quant | exec:update_counts --check -> 0 |
✅ |
On a 32-file labelled clean-UI corpus, 18 of the 19 shipped design-slop rules produced zero false positives; the nineteenth (slop-c6-lock-colour, catalog C6) fired on 4 of the 32 files and was demoted to judgment-only rather than tuned. |
quant | internal/bench/corpora/design-slop-fp-PREREG.md#The ceiling, declared before the run |
✅ |
| On a weak host (claude-haiku-4-5) the package produces a significant, placebo-controlled discipline lift on scope/downstream traps; on a strong host the same measurement is a published null — the package transplants discipline a weak model lacks, not model intelligence. | quant | docs/benchmark.md#weak-host-specific |
✅ |
| The non-coding domain skills (finance/founder/ops/content) are forged on TS/PHP and labeled unvalidated until they pass a sourced domain-truth fixture; no public prose implies proven domain correctness, and the validated count is CI-ratcheted. | qual | exec:domain_soundness_status --check -> 0 |
✅ |
The validated non-coding domain-skill count is pinned and CI-ratcheted at a maintainer-set floor (9 of 20 default-surface skills carry a sourced evals/domain-truth.json fixture at pin time, 2026-07-11 — 5 deterministic, keys from cited formulas; 4 rubric, criteria matching a named external practice); the floor only rises via a maintainer --write-floor after a new sourced fixture lands. |
quant | exec:domain_soundness_status -> 0 |
✅ |
| On the READ-ONLY FAN-OUT slice family, tier-downshifted subagent dispatch (lite/haiku vs session-tier-proxy sonnet) nets a ≥30% USD-weighted token-cost reduction at held quality — measured 2026-07-08 (n=10 paired live dispatches, 20 telemetry lines): 10/10 exact-match on BOTH arms, 29.4% fewer raw tokens, 76.5% USD-weighted cost reduction at the 3x haiku↔sonnet price ratio. FAMILY-SCOPED — the mechanical-edit family is unmeasured and its downshift (incl. the deferred tier downgrades of existing units) stays gated. Negative control held: an open-ended synthesis/unknown slice never resolves below the session tier (inferSliceTier → medium/inherit, never lite). | quant | internal/bench/routing-downshift/results-2026-07-08.md#FAMILY-SCOPED PROVE |
✅ |
The text-layer-only boundary above is machine-enforced, not asserted: a scope-guard test fails if any corpus entry declares a non-text layer, and that guard is itself falsified by a test that splices in a layer: file PNG-metadata fixture and requires the guard to fail. The corpus is sha256-frozen and was committed BEFORE any detector existed, so no detector was tuned against it. |
qual | tests/scripts/encoding_corpus.test.ts#FAILS when a deliberately out-of-scope fixture is added |
✅ |
| The retrieval sanitize floor covers the TEXT layer only, by construction — never file or network channels (image / audio / PDF / DNS / TCP / file-metadata steganography) and never semantic evasion (word choice, phrasing, garden-path constructions, word-order permutation). On the frozen 653-entry corpus it strips or flags 99.00% of in-scope positives (100.00% on the unambiguous zero-width / bidi / variation-selector classes) at a 0.00% false-positive rate over 353 real in-repo negatives, with zero added model spend and 0.018 ms p95 per message. Exactly ONE of the seven added channels removes bytes; the other six report and pass the text through unchanged — this is NOT a claim to block steganography. | quant | agents/evidence/reports/encoding-floor-measurement.md#Selected branch: ADOPT |
✅ |
The share of governed rules carrying a backstop that fails a CI build is RESOLVED, not declared, and is published in exactly ONE place — docs/proof.md § 4b, projected from check_enforcement_coverage, which prints its denominator together with the frame that produced it (internal/reports/enforcement-coverage.json is that same output on disk). No figure is restated in this entry, deliberately, and check_enforcement_denominator reds when one appears in a published doc the resolver did not generate: until 2026-08-23 this tree carried FIVE different numbers for the one property, and every previous correction fixed a figure while leaving the plurality intact — a number that is right today is how it came back each time. Resolution means a declared validator: counts only when the script exists AND a GITHUB WORKFLOW reaches it (transitively, so a sub-check under a wired umbrella counts), while a hook registered fail_closed: false resolves to observer, never validator. That distinction is not cosmetic: it once let “named in a taskfile” read as “fails the build” while NO workflow invoked task ci, ci-strict, or ci-fast, so most of the counted validators only ran when a human typed the command. Splitting validator (CI runs it) from validator-local (only a human does) cut the honest figure to roughly a third at the time; wiring the taskfile-only gates into rule-backstops.yml restored it, this time meaning what the headline says, and local_only is ratcheted at zero so a gate cannot drop back out silently. Wiring them also surfaced that five were failing invisibly, 37 findings deep; those are cleared and the baseline in rule-backstop-debt.json stands at zero, so the ratchet enforces rather than tolerates. Roughly two thirds of the 37 were never violations — the gates were misreading allowances their own rules already grant (license-required attribution, multi-stack peer examples), which is the same class of defect one level down. The ratio has not risen across five releases, and the reason is worth stating rather than reading as stagnation: the ratchet prevents regression, it does not raise the level. An undeclared rule counts as uncovered, never excluded — an honest recorded gap beats a false claim of coverage — including the two scale/history pack rules that ship enforced by lint_persistence in consumer CI, which this resolver, scoped to THIS repo’s workflows, correctly does not count. |
quant | exec:check_enforcement_coverage --check -> 0 |
✅ |
| Exactly one enforcement denominator is quotable, it names the frame that produced it, and no published doc restates it by hand. | quant | exec:check_enforcement_denominator -> 0 |
✅ |
| The lift-carrying essential cut (kernel + downstream-changes) keeps a significant weak-host discipline lift at a fraction of the full load’s tokens, and the lift is FAMILY- and HOST-SCOPED — measured on three hosts: claude-haiku-4-5 (weak) shows the family-scoped lift (trapE 0.533→1.000, 7/7 discordant, corpus cost 1.71x); claude-sonnet-4-6 (strong) is a ceiling null; gpt-5-mini (non-Claude weak, codex prompt-prepend surface) FAILED replication with headroom (corpus Δ=+0.024 p=0.70, capability trend n.s. — no harm claimed, injection-surface confound documented). Therefore discipline_profile: auto enables the lift only where measured (vendor-granular unknown_defaults). Non-claims — the balanced router profile was removed after a NULL measurement (p=0.81, n=24); no full-tier recommendation exists; no cross-vendor lift is claimed. | quant | docs/benchmark.md#REPLICATION FAILED |
✅ |
| Behavioral-eval coverage is measured per tier and CI-ratcheted so it can only rise; the current coverage and its gap are published, never implied as “264 evaluated skills”. | qual | exec:skill_eval_coverage --check -> 0 |
✅ |
| A credential-free prescription layer reads content the host’s own web tools cannot fetch at all — Reddit thread text (Atom feeds) AND comment ranking plus reply nesting (server-rendered HTML), and a single named tweet (the platform’s own oEmbed endpoint). Measured 2026-07-25 from a residential network against a pre-registered 6-task-per-channel set with a native control and thresholds frozen before the run: reddit tier 1 6/6, reddit tier 2 6/6, twitter-oembed 6/6, native 0/6 on both Reddit tiers. Zero credentials, zero resident processes, zero auto-installs. Scope bounds that travel WITH the claim: (a) the twitter gap is narrower than 6/6 suggests — native also passed 2 of those 6 via third-party mirrors, so the channel earns its place only on tweets nothing mirrors; (b) reddit tier 2 is on an announced closing path and ships with a kill-switch keyed on an OBSERVED login wall; (c) youtube-transcripts is PARKED, not shipped — its backend is human-installed by contract and was never exercised; (d) residential network is load-bearing, and CI is explicitly not a bench environment. | quant | docs/benchmark.md#ship-gated-reach |
✅ |
Four adjacent governance properties were closed as regression tests rather than phases, on the expectation that each was already true. One of the four was. (a) enforcement never branches on a base-model refusal string — holds; the single module compiling refusal regexes only ever escalates, and no refusal branch reaches an allow decision. (b) a capability gate resolves only from trusted config — VIOLATED: the runtime dispatcher returned ready for a skill whose own frontmatter declared safety_mode: strict and granted itself 2 tools absent from the 2-entry registry, while the validator implementing that allowlist had zero production callers; now wired. (c) caller-agnosticism — holds: 0 caller-identity inputs reach a gate verdict, and the 3 platforms that carry the blocking slot are pinned so none silently loses a concern. (d) constraint monotonicity — holds: 0 blocking gates read persisted state, with 1 advisory anti-nag exception named in the test. All 7 inverted properties produced a failing test. |
quant | internal/bench/reports/governance-invariants.json#adjacent_properties |
✅ |
The council aggregation cannot be steered against a refusal. Pre-registered spike S0.1 measured that it WAS classification-steerable — w_total counted only members whose stance line parsed, so a refusal phrased as prose left the quorum and made consensus easier: steering margin 0.6667 (margin -0.25 parsed vs +0.4167 unparsed) with the outcome flipping from no-consensus to Adopt. High severity because the direction was the dangerous one. Fixed in the same change: a member who responded counts toward the quorum whether or not its stance parsed, and the post-fix steering margin is exactly 0. The divergence signal is an observation and is asserted never to reach the scoring path. No observed attack prompted this; the expected outcome was a null and it was pre-registered as such before the run. |
quant | internal/bench/reports/governance-invariants.json#s0_1_aggregation_steerability |
✅ |
Pre-registered spike S0.2 measured that both of this package’s fail-closed gates judged the shape of ONE action rather than the effect, so a sequence whose every step they allowed composed into the outcome they exist to prevent — leak count 2 of 2 gated outcomes, with all 4 single-step controls blocking correctly. Both were moved to the effect they govern. The positive control is check_secret_leak, which scopes to the cumulative diff against a base ref and therefore returned null before any fix — decomposition gains nothing against an effect-scoped gate. Scope bound, published rather than swept: mv / chmod / rm against .git/hooks/* still reach the first outcome and are asserted as an open gap, because recognising them would make a fail-closed guard a shell sandbox. Decomposition in this package is model-carried — there is no executable subagent dispatcher — so the measured layer is the PreToolUse hook layer; the 2026-08-02 council dissent on whether that is a faithful discharge is recorded in the spike header. |
quant | internal/bench/reports/governance-invariants.json#s0_2_decomposition_laundering |
✅ |
| Pre-registered spike S0.3 measured that a stated uncertainty, hedge, or provenance marker survives this package’s telegraph condenser into the audited text — 3 marker classes, 10 fixture cases, marker-loss count 0, negation count preserved. Honest null, published with the spike wired as the regression test. The first run’s two “failures” were fixture defects (carriers written as phrases containing an article the condenser is documented to drop) and are recorded as an unmet premise rather than edited away. Scope bound: this protects a marker the agent DID emit; it cannot make an agent state an uncertainty it never stated. | quant | internal/bench/reports/governance-invariants.json#s0_3_marker_survival |
✅ |
Hook dispatch runs as one precompiled node process with all concerns in-process, and the CI latency gate measures the REAL invocation path — the exact command hooks.json installs (bash wrapper + install-shape probes + dispatcher), via bench_hook_latency --gate --via-cli, whose per-event commands come from the same generator that writes hooks.json — not the bare bundle the pre-repair gate measured. The repair (road-to-hook-latency-repair) is pinned before/after on one machine in the committed baseline history: pre-fix CLI path pre_tool_use p95 164 ms → post-fix bundle-direct path p95 84 ms (darwin dev, warm cache, n=50 each; the pre-fix path measured ~450–500 ms/event on a 1-vCPU container). Budget re-derived 2026-08-19: pre_tool_use p95 175 ms (was 150), any event 250 ms on GitHub-hosted CI runners, both BLOCKING. The raise is measured rather than bent around a red check — 150 sat INSIDE the observed legitimate distribution (green main runs 115, 120, 146, 150, 150 ms; reds 151, 152, 152, 154 ms; a comment-only PR run at 157 ms) and failed 5 of 33 tests.yml runs on main over 2026-08-16..18, i.e. 15%, with hand re-runs as the standing workaround. The budget file records that distribution, the failure rate, the cited derivation rule and a revisit-if that routes the NEXT breach to the cause rather than to a third raise: a GREEN run whose p50 rises above 160 ms is the cost growing, not the runner. Down from ~1.6 s p50 on the retired CLI-to-bash-to-tsx per-concern-respawn chain. |
quant | docs/hook-latency.json#invocation_path |
✅ |
23 host agents are detected and inventoried; 20 receive a written config surface (18 projection + 1 plugin + 1 bundle target) and 3 are export-only (aider, zed, jetbrains). The count is enforced, not asserted — knownToolIds() is pinned at 23 by a test whose assertion literal IS the number, and src/config/surface-matrix.yml is held in set-equality with the installer’s own user-scope path map by lint_surface_matrix, so a host added to one and not the other fails the build. This entry stood unbacked while naming its own unblocking condition (“once surface-matrix.yml exists, bind the count to that file and flip”); the condition was met and nothing fired, so the shipped figure stayed “7+” — understating real coverage by 3x. |
quant | exec:vitest run tests/install/toolDetection.test.ts -> 0 |
✅ |
On the fixture corpus (n = 20 before/after pairs, 17 length-controlled within ±25%), the humanizer pass removes every mechanically detected AI-writing tell (mean hard hits 0.9 → 0, cluster score 48.06 → 0 per 500 words, dash density 8.76 → 0), and a blind judge (claude-sonnet-4-5, deterministic per-pair A/B seed) preferred the humanized text in 16/16 of the pairs that were length-controlled when the judged run fired. Two figures moved on 2026-09-07 and neither move is a re-measurement of the same thing: the before-side cluster score went 53.97 → 48.06 under two scoring changes pulling in opposite directions (thirteen tell families landed, raising it; a pattern used uniformly through a document is now charged once instead of N times, lowering it further), and the length-controlled count rose 16 → 17 because a repair to the scanner’s quoted-span strip changed word counts. The fall is not a recall loss and the rise was not a recall gain — the scoring rule changed, so the aggregate is not comparable across the boundary, and per-family recall is recorded in internal/bench/corpora/prose-tells-epochs.md instead. The seventeenth pair has not been judged; the 16/16 figure is the sixteen that were. Every objective figure is also reported per tune/holdout split in the report, so a number cannot be read as held-out when it was not. Scope note — the “before” fixtures were deliberately tell-seeded, so this measures seeded-tell removal on a self-constructed corpus, NOT real-draft improvement. Fixture-only evaluation is the declared scope of the current roadmap phase (road-to-measured-prose-tells, Phase 3): collection of real drafts, and of metrics derived from real drafts, was declined for this round because either would create a new retention practice beyond the fixture-only data-handling floor this package currently records. That is a decision about this round and not a permanent boundary — future authorization is neither granted nor refused here, and the owner-facing question stays exactly where it was. Real-draft lift therefore remains unmeasured, and this claim will not widen beyond “on the fixture corpus” until it is both authorized and measured. |
quant | internal/bench/reports/humanizer-v1.md#prefers the humanized text |
✅ |
The injection-scan post_tool_use detector measures 99.00% recall (297/300) and a 0.85% false-positive rate (3/353) over the frozen corpus at internal/bench/corpora/encoding-channels/. Both figures are properties of THAT corpus and not of tool output in the wild — the evidence report says so itself: web fetches and MCP responses are a different distribution, so the 0.85% predicts nothing about a repository whose tool output is largely prose about prompt injection. Distinct from encoding-floor-text-layer-only, which measures the stripping pipeline at a 0.00% false-positive rate over a different scope; the number here belongs to the reporting half, which had never been measured separately before. The detector ships enabled: false and warn-only, so neither figure is a claim that anything is blocked. |
quant | agents/evidence/reports/injection-detector-wiring.md#The numbers |
✅ |
| A fresh registry install of this package carries zero high/critical npm-audit findings on the runtime dependency tree (0 vulnerabilities total at last verification), gated on every PR and every release PR. | quant | .github/workflows/release-validation.yml#npm audit --omit=dev --audit-level=high |
✅ |
The judge-skill family (judge-bug-hunter, judge-code-quality, judge-security-auditor, judge-synthesis, judge-test-coverage, and the /review-changes dispatcher) implements the LLM-as-a-judge pattern — a specialized model scoring another model’s output against a rubric — and names position bias and self-consistency as its known failure modes. SCOPE: the pointer backs the pattern and its named failure modes, which is what the cited work establishes. It does NOT back any claim about how this repo mitigates them; a statement about the dispatcher’s own behavior needs a repo pointer, not a paper. |
qual | https://arxiv.org/abs/2306.05685 (2026-08-11) |
✅ |
| Word-boundary-anchored keyword matching reduced unintended rule activations by 12.5% (495 → 433) over the 302-prompt matrix-derived corpus with zero intended positives lost. Disclosure — the derived corpus was co-edited in the same change (6 German positives re-authored to standalone tokens; verb-inflection recall is a documented accepted cost). The circularity is broken by an independent replay over 49 UN-edited real-corpus prompts: recall 15/17 in BOTH arms (zero labels lost to anchoring), unintended activations 110 → 99 (−10.0%). | quant | agents/evidence/analysis/anchoring-independent-replay-2026-08.md |
✅ |
NONE of the backed ledger claims are machine-re-verifiable today — every evidence pointer is checked for existence (file present, substring present, URL carries a date), never for truth, so a claim pointing at a stale artifact stays backed indefinitely. A measured minority COULD carry a re-executing exec: form, clearing the >= 10 pp threshold that was pre-registered before the count was taken, which is why that form is scheduled rather than assumed. The rest cannot: paid or stochastic benchmark runs no CI job can re-derive, and prose contracts. Exact counts live in the evidence file and are NOT restated here on purpose — this entry hard-coded its denominator twice and drifted within a day both times (25 when the ledger held 26, then 26 when it held 27) while CI stayed green, because the pointer resolved. check_claims now compares the stored denominator against the live ledger and fails on divergence. A number a human retypes on every ledger edit will drift; the fix was to stop retyping it. |
quant | internal/reports/exec-evidence-feasibility.json#"exec_feasible" |
✅ |
A hand-rolled, dependency-free BM25 + trigram lexical index resolves the “recalls but does not rank” gap: on the retrieval-precision corpus (9 keyword-overlapping-confuser tasks) it drives the mean top tie-set from 3.333 (the _score bucket scorer) to 1.0 — every needed decision uniquely top-ranked — with precision@1 and precision@5 unchanged at 1.0. Method: deterministic, model-free re-ranking of the SAME retrieved entry set; both scorers measured over the identical store via measure_lexical_ranking.ts. Cross-artifact note (2026-07-25): this baseline (3.333) is the value in THIS claim’s own artifact and is cited correctly, but the retrieval-precision artifact records 4.111 for the same scorer on the same corpus — two bench scripts disagree. The lift direction (ties collapse to 1.0) holds under either baseline; the discrepancy itself is unresolved and recorded rather than smoothed. |
quant | exec:measure_lexical_ranking -> 0 |
✅ |
| Registering an MCP server with this package costs standing context on every session, and the kernel server’s 25-tool surface costs 4,603 tokens of it while the two-tool lite surface is capped at 600. | quant | agents/evidence/metrics/mcp-tool-standing-cost.jsonl#tool_search_threshold |
✅ |
| On a 12-fixture option-decision corpus (3 arms × 2 providers, blind rubric judge claude-opus-4-8, pre-registered hypotheses), famous-figure identity framing added nothing beyond the underlying method text (method 5.04 vs figure 4.88, Δ=0.17, sign-test p=0.607), and provider diversity moved judged quality ~15× more than persona identity (provider Δ=2.58 vs identity Δ=0.17); the whole persona layer lifted only +0.08 over bare prompts. Honest null — persona panel-mode stays CUT, evidence-closed. | quant | internal/bench/reports/persona-placebo.json#honest-null |
✅ |
| No readable external-source attribution exists in the tracked tree — not in a file’s content and not in a file’s path — and every new occurrence fails CI rather than being discovered later. | qual | exec:check_no_external_sources -> 0 |
✅ |
| We publish our own measured null results and retire or constrain features when the evidence does not support them. Deliberately falsifiable — every published null links the run that produced it; find one that does not resolve and this line updates. | qual | docs/benchmark.md#honest |
✅ |
Every artifact count this package publishes about itself — the six badge integers for skills, rules, commands, guidelines, personas and advisors — is re-derived from the tree by one canonical counter rather than hand-typed, and a drift of even one fails CI. Measured 2026-09-28 via that counter: skills 299, rules 121, commands 203 (recursive; 60 top-level), guidelines 121, personas 29 (README excluded), advisors 5. The counting BASIS differs per noun and is not inferable from the directory the badge links to — commands counts recursively while the linked directory holds 60 top-level files, and rules counts the 121 source rules while the linked projection holds 120 because one dormant rule is not projected. Both bases are stated next to the badge block, because an undeclared basis is not a wrong number but an unreadable one. |
quant | exec:update_counts --check -> 0 |
✅ |
On the live end-to-end retrieval benchmark (9 tasks × 3 arms × 3 seeds on claude-haiku, fixture store with keyword-overlapping confusers), the memory retrieval substrate scored precision@5 = 100% (9/9) with 100% poisoned-entry rejection, and the retrieval-on arm passed 27/27 model-scored tasks vs 2/27 with retrieval off and 4/27 with a placebo injection. Known limit stays published: mean tie-set 4.111 means top-k ties break by store order, not relevance (the ADR-116/FTS5 signal). SOLE RECORD for this artifact as of 2026-07-25: a second entry (second-brain-retrieval-precision) described the same measurement from the precision angle and had drifted to 5/27, 5/27 and tie-set 3.3 — figures absent from the shared artifact. Two entries over one artifact is what allowed them to disagree while both resolved, so the pair was folded into this one. |
quant | internal/bench/reports/second-brain-retrieval.json#retrieval-on |
✅ |
| 122 governed rules. | quant | exec:check_artefact_count_messaging -> 0 |
✅ |
Scope de-duplication of the rule projection removes 38.0% of the median cold-start payload (87,677 of 230,556 tokens) against a pre-registered 15% bar. CONDITION, inseparable from the number: that figure is FIXTURE-MEASURED on byte-identical global and project projections, and it is currently unreachable for production installs — the installer stamps ownership metadata (package: / source_path:) onto every installed rule unconditionally, so the two scopes are produced by two writers with deliberately different output, the byte-identity gate correctly refuses to dedup, and the recipient set is empty. The mechanism works; the saving is not realised by any consumer today. |
quant | agents/settings/contexts/cache-economy-refusals.md#Honest null — scope de-duplication is measured but |
✅ |
| On a deterministic multi-session recall corpus, the memory substrate produces a measured, placebo-controlled recall lift — memory-on 27/27 vs no-memory 10/27 and vs equal-byte placebo 9/27 (claude-haiku-4-5, n=9 tasks x 3 seeds, sign test p=0.031 for BOTH pairings). Scoped honestly: this is the context-value upper bound (perfect retrieval on a one-fact-per-task corpus), not retrieval precision under a large store. | quant | internal/bench/reports/second-brain-delta.json |
✅ |
sequential-thinking applies chain-of-thought decomposition with two constraints the original formulation does not carry — a cap on the number of thoughts and a mandatory validation step — specifically to bound the unbounded-expansion failure mode. |
qual | https://arxiv.org/abs/2201.11903 (2026-08-11) |
✅ |
Every artifact the package ships — source AND the condensed projection that reaches consumers — is machine-scanned in CI for hidden-Unicode, mixed-script-confusable, and instruction-smuggling payloads (the rules-file-backdoor class); a finding blocks the release before npm publish, not just the merge. |
qual | exec:lint_agent_security -> 0 |
✅ |
Skill self-selection is measured, not assumed, and over this package’s own transcript store it is ZERO — not near zero. report_skill_activation over 30 sessions and 11,338 assistant turns records 0 Skill invocations and 0 of 299 distinct skills (measured 2026-09-06). Every figure in this sentence, and the date, is read out of agents/evidence/metrics/skill-activation-census.json, which the census writes and refuses to write from an empty store; check_skill_activation_claim fails when the record and this sentence disagree, so the published number cannot outlive its measurement silently. The 299 split three ways and the zero means a different thing in each: 12 declare a machine-matchable trigger key in frontmatter, 100 carry an evals/triggers.json corpus, 2 do both, and the remaining 189 are reachable only by a human naming them. For the 12 the zero was published as a defect — a matchable declaration exists and nothing acted on it — and ADR-263 SUPERSEDES that reading in those words: the 12 are unvalidated metadata with no host contract to bind to, so nothing was owed and nothing failed; they are kept, not deleted, as the thing a future host contract would bind to. For the 189 the zero is the design, because no automatic selection is claimed for them and invocation there is a reader opening a file. For the 100 the zero licenses nothing in either direction: evals/triggers.json is a TEST fixture read by check_routing_coverage / lint_skill_trigger_corpus / check_trigger_evals, no host reads it at routing time, and the two surfaces share the word “trigger” and are unrelated. A 14-file eval-corpus wave taking coverage 100 → 114 was authored, measured against, and then reverted on the same branch for an unrelated governance reason; the census read 0 in both states. FRAMING, RECORDED 2026-09-07 SO A LATER REVIEW MEETS AN ANSWER RATHER THAN RE-DERIVING THE ARGUMENT: the surface is left exactly as it is and the zero is published with its reason. Two substantive alternatives were on the table and NEITHER was taken — building a host-side selection path for the population that declares a machine-matchable trigger key, or declaring the human-named remainder reference material by design. Each of those changes what the package claims to be for its consumers, which is an owner-reserved commitment under decision-revisit-gate — and on 2026-09-08 that commitment was made: ADR-263 LOCKS the second, in the corrected form “by observation” rather than “by design”, and REJECTS the first, on the ground that there is no host contract to build against and that shipping a documented routing capability which still yields zero would convert an implementation cost into a credibility cost. So neither alternative is open any longer, and the sentence that said both were is corrected here rather than left standing. What ADR-263 did NOT do is still worth naming, and the boundary is narrower than a first reading suggests: no selection mechanism was built, and the numeric split above — 12 / 100 / 2 / 189 — is untouched. It DID reclassify the 12 semantically, from a population owed a selection path to unvalidated metadata with no host contract, which is the supersession recorded two sentences up; what it did not do is move a skill from one of the four counts into another — “by observation” is deliberately not “by design”, because design intent is an ownership claim this repository cannot make about hosts it does not own. Its own falsifier, stated so it can fail: the reason for accepting the zero as-is is that a single-machine transcript store cannot separate an unbuilt selection path from an unused one, so the moment a second store — a consumer install, another host, a CI-visible corpus — becomes readable, the reason stops holding and the choice returns. |
quant | exec:check_skill_activation_claim -> 0 |
✅ |
| 299 skills. | quant | exec:check_artefact_count_messaging -> 0 |
✅ |
skill-improvement-pipeline converts post-task outcomes into durable written lessons rather than in-context retries, following the Reflexion formulation. |
qual | https://arxiv.org/abs/2303.11366 (2026-08-11) |
✅ |
Every cross-skill SKILL.md link in the authored skill corpus resolves on disk, and the census behind that statement is derived by the same collector the gate scans with. |
quant | agents/evidence/metrics/skill-link-census.json#"dead_links": [] |
✅ |
Whether serving low-priority skills over MCP instead of listing them natively improves skill selection (H1) is NOT established, and projection.mode: tiered therefore stays opt-in. |
quant | agents/evidence/analysis/skill-tiering-matrix-arm.md#The question this arm cannot answer |
✅ |
On a default Claude Code install projection.mode: tiered costs MORE standing context than legacy-all — 2,259 tokens against 1,956 — because the host already caps description delivery at roughly the Tier A set; the 82% saving exists only against a 100%-delivery counterfactual. |
quant | agents/evidence/analysis/skill-tiering-matrix-arm.md#H2 |
✅ |
Removes only its own keys from a shared host config, and leaves a neighbour tool’s entries byte-identical — including inside a shared hooks.<event> array, where a key this package merely appends to is identified by its own command signature rather than claimed wholesale. |
qual | exec:vitest run tests/lib/json_pointers.test.ts tests/lib/host_hook_merge.test.ts -> 0 |
✅ |
| On the pre-registered 12-fixture defect-finding corpus (three arms, deterministic file-level recall, codex reviewer gpt-5.5), the cross-model team-review arm produced NO recall lift over single-model self-review — all three arms recalled every planted defect (Δ = 0, H1 not met). Honest null, ceiling-limited (recall 1.00 everywhere: the seeded defects are too obvious to discriminate the arms on recall); the only non-null signal is a single self-review false positive on the controversial-but-correct control vs 0 for team/council. No cross-model quality/lift claim binds; team mode stays workflow-value-only. Re-open: a judge-survivable-subtlety corpus or a new model generation. | quant | internal/bench/reports/defect-finding.json#honest_null |
✅ |
SHIPPED FOR CLAUDE CODE ONLY, and the per-host scope is the whole claim — a lean_projection.mode: delivery projection plus the rule-inject hook concern delivers rule bodies on a trigger match. Re-measured 2026-09-08 over the widened corpus: 591/591 byte-equal deliveries with 0 unequal, 101/101 labelled rules reachable with open_files honoured, 0 false fires over 212 labelled near-misses, and $0.7047 vs $4.0497 per 50-turn x 5-spawn session at sonnet rates. THE EVIDENCE POINTER MOVED ON 2026-09-08 (R2 finding 7 on PR #1923) and the old one was wrong: this row credited thin-inject-2026-08-23.md#endpoints with the re-measured figures, and that report carries 579/579 over a 94-rule corpus — check_claims only verifies that a pointer RESOLVES, so the row read backed while its evidence predated the measurement. It now cites thin-inject-2026-09-08.md, written from the run it reports. TWO FURTHER CORRECTIONS in the same review, both narrowing what this row may be read to say. The delivery count fell 616 -> 591 and the price 0.7167 -> 0.7047 because the concern’s CAP_BYTES was lowered 20,480 -> 16,384 B onto the user_prompt_submit slot row it had exceeded by 25 % (finding 3), so more bodies are withheld on large fires: fires truncated 33 -> 45, bodies withheld 63 -> 88. That is a delivery REDUCTION and not an improvement, and endpoint (a)’s bar is unequal == 0, unaffected by the count. And the recall figure now carries a second reading: 101/101 is measured with open_files HONOURED, which is the reach of the MECHANISM and not of the shipped binding — rule-inject is bound on user_prompt_submit + pre_compact only and that slot never populates openFiles — so the endpoint also publishes the shipped reach, 99/102 as of 2026-10-02 (98/101 at the 2026-09-08 re-measurement; the corpus gained one labelled rule), and names the three rules that differ (design-review-after-ui-write, source-of-truth, ui-audit-gate, all path-only and all kept full-bodied by the always-eager residue, so no THINNED rule is unreachable). On a clean consumer-shaped root the tree Claude Code loads falls 99,598 -> 24,166 chars/4 tokens at 114 files on BOTH sides, while .cursor/rules and .clinerules stay byte-identical at 121,242 — asserted per PR by check_host_tree_parity, not merely observed here. WHAT CHANGED ON 2026-09-07, stated because this row previously said the opposite and a reader may have quoted it: the mode was MEASURED-BUT-NOT-SHIPPED with the concern default-OFF and the activation charge unpaid; ADR-267 makes delivery + hosts: [claude-code] the shipped template default, the charge is paid (the user_prompt_submit slot sum moved 4,096 -> 16,384, the measured p90 gate-open fire size rounded up to 512), and the pre_tool_use binding was REMOVED rather than kept, so that slot’s 2,048 cap is untouched. The earlier figures (579/579, 94/94, 194 near-misses, $0.7285 vs $4.0335, 120,582 -> 18,573 exact-BPE) were measured on a 94-rule corpus and are superseded rather than deleted: the corpora differ, so no delta may be computed between the two sets. Latency measured flipped: pre_tool_use p95 125 ms and user_prompt_submit gate-open p95 81 ms, both inside the 175 ms p95_ci budget. What these endpoints license is delivery equivalence and cost, and nothing wider — the paired-judging instrument is closed by ADR-202 at inter-evaluator Cohen’s kappa 0.472 against a registered 0.800 floor, and neither the flip nor this re-measurement reopens it. Method: model_rule_injection --endpoints over the frozen tests/eval/routing-matrix corpus, pre-registered in internal/bench/thin-inject-PREREG.md before the report artifact existed, with --selftest requiring each endpoint to reject a planted defect first. |
quant | internal/bench/reports/thin-inject-2026-09-08.md#endpoints |
✅ |
| Over the frozen ui-conformance fixture, the probe reports all four planted behavioral defects and raises zero findings for the one declared deviation, against a pre-registered screenshot arm that catches two of the same four and raises one. | quant | tests/scripts/ui_conformance_probe.test.ts#finds four of four planted defects |
✅ |
On the A3 orchestration eval (deterministic verify, measured token deltas), the production-validator subagent returned the correct verdict on both fixtures — NOT READY with the exact file:line citation on a planted hollow implementation, READY with zero spurious findings on the clean control — while consuming ~45k fewer tokens than the inline-host baseline on each task. Scope: two planted fixtures on a Claude Code host, not a broad hit-rate. |
quant | internal/bench/orchestration/pv-a3-results.md#token_delta_provenance: measured |
✅ |
61 backed claim(s) — all evidence pointers resolve in CI.
How many of those re-derive themselves
Section titled “How many of those re-derive themselves”18 of 61 backed claims carry exec: evidence —
CI re-runs the command and compares its exit code to the claim, so a
stale one turns the build red. The other
43 rest on a pointer: CI checks that the artefact
exists and contains what it should, which cannot distinguish a live
claim from one whose producer nobody has run in months.
That residue is not an oversight, and it is listed rather than rounded away. A claim can only re-derive itself when a deterministic command carries the verdict in its exit code. These cannot, for the reasons given:
| Claim | Why it cannot re-execute |
|---|---|
| The two surfaces the provider-lifecycle contract obliges to agree on an adapter tier — the adapter header and | prose or contract artefact — no exit code carries the verdict |
adversarial-review structures its critique as branching exploration with explicit pruning rather than a sing |
external cite — CI does not fetch the network |
The .augment-plugin/ manifest version is the package version, not an independent plugin-API version, and eve |
prose or contract artefact — no exit code carries the verdict |
analysis-autonomous-mode runs iterative self-critique between steps rather than only at the end, following t |
external cite — CI does not fetch the network |
bug-analyzer verifies each candidate root cause against a concrete trigger before reporting it, following th |
external cite — CI does not fetch the network |
| The release process is documented as an inheritable runbook + succession doc, and the project’s bus-factor (tr | prose or contract artefact — no exit code carries the verdict |
| On the pre-registered 2-arm retrieval benchmark (18 hand-verified code-structure questions across 3 real repos | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
On the first post-fix behavior-conformance measurement (round 5, /analyze:conformance --limit 30, run 2026-0 |
benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| A MEASURED-BUT-NOT-SHIPPED experiment — the thin rule projection reduced eager rule load 78,513 → 13,881 GPT-t | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| The first cross-vendor parity pass (5 orchestration-corpus tasks × 2 vendors × 3 repeats, identical prompts vi | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| On a 32-file labelled clean-UI corpus, 18 of the 19 shipped design-slop rules produced zero false positives; t | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| On a weak host (claude-haiku-4-5) the package produces a significant, placebo-controlled discipline lift on sc | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| On the READ-ONLY FAN-OUT slice family, tier-downshifted subagent dispatch (lite/haiku vs session-tier-proxy so | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| The text-layer-only boundary above is machine-enforced, not asserted: a scope-guard test fails if any corpus e | prose or contract artefact — no exit code carries the verdict |
| The retrieval sanitize floor covers the TEXT layer only, by construction — never file or network channels (ima | prose or contract artefact — no exit code carries the verdict |
| The lift-carrying essential cut (kernel + downstream-changes) keeps a significant weak-host discipline lift at | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| A credential-free prescription layer reads content the host’s own web tools cannot fetch at all — Reddit threa | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| Four adjacent governance properties were closed as regression tests rather than phases, on the expectation tha | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| The council aggregation cannot be steered against a refusal. Pre-registered spike S0.1 measured that it WAS cl | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| Pre-registered spike S0.2 measured that both of this package’s fail-closed gates judged the shape of ONE actio | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| Pre-registered spike S0.3 measured that a stated uncertainty, hedge, or provenance marker survives this packag | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| Hook dispatch runs as one precompiled node process with all concerns in-process, and the CI latency gate measu | prose or contract artefact — no exit code carries the verdict |
| On the fixture corpus (n = 20 before/after pairs, 17 length-controlled within ±25%), the humanizer pass remove | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
The injection-scan post_tool_use detector measures 99.00% recall (297/300) and a 0.85% false-positive rate ( |
prose or contract artefact — no exit code carries the verdict |
| A fresh registry install of this package carries zero high/critical npm-audit findings on the runtime dependen | prose or contract artefact — no exit code carries the verdict |
The judge-skill family (judge-bug-hunter, judge-code-quality, judge-security-auditor, judge-synthesis, |
external cite — CI does not fetch the network |
| Word-boundary-anchored keyword matching reduced unintended rule activations by 12.5% (495 → 433) over the 302- | prose or contract artefact — no exit code carries the verdict |
| NONE of the backed ledger claims are machine-re-verifiable today — every evidence pointer is checked for exist | prose or contract artefact — no exit code carries the verdict |
| Registering an MCP server with this package costs standing context on every session, and the kernel server’s 2 | prose or contract artefact — no exit code carries the verdict |
| On a 12-fixture option-decision corpus (3 arms × 2 providers, blind rubric judge claude-opus-4-8, pre-register | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| We publish our own measured null results and retire or constrain features when the evidence does not support t | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| On the live end-to-end retrieval benchmark (9 tasks × 3 arms × 3 seeds on claude-haiku, fixture store with key | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| Scope de-duplication of the rule projection removes 38.0% of the median cold-start payload (87,677 of 230,556 | prose or contract artefact — no exit code carries the verdict |
| On a deterministic multi-session recall corpus, the memory substrate produces a measured, placebo-controlled r | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
sequential-thinking applies chain-of-thought decomposition with two constraints the original formulation doe |
external cite — CI does not fetch the network |
skill-improvement-pipeline converts post-task outcomes into durable written lessons rather than in-context r |
external cite — CI does not fetch the network |
Every cross-skill SKILL.md link in the authored skill corpus resolves on disk, and the census behind that st |
prose or contract artefact — no exit code carries the verdict |
| Whether serving low-priority skills over MCP instead of listing them natively improves skill selection (H1) is | prose or contract artefact — no exit code carries the verdict |
On a default Claude Code install projection.mode: tiered costs MORE standing context than legacy-all — 2,2 |
prose or contract artefact — no exit code carries the verdict |
| On the pre-registered 12-fixture defect-finding corpus (three arms, deterministic file-level recall, codex rev | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
SHIPPED FOR CLAUDE CODE ONLY, and the per-host scope is the whole claim — a lean_projection.mode: delivery p |
benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| Over the frozen ui-conformance fixture, the probe reports all four planted behavioral defects and raises zero | prose or contract artefact — no exit code carries the verdict |
| On the A3 orchestration eval (deterministic verify, measured token deltas), the production-validator subagent | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
Artefact counts in public prose (skills, commands, governed rules,
guidelines, personas) are generated from source and CI-drift-checked:
update_counts.ts writes the numbers, check_artefact_count_messaging.ts
fails the build on any count-shaped prose mention that drifts from the
source count — or on two different numbers for the same artefact kind.
We also publish our debt: 35 claim(s) are logged as
unbacked inventory in the ledger — not yet bound, and therefore not
allowed to carry a marker in public prose. Hiding them would be the
opposite of the point.
And our nulls: 7 claim(s) are resolved-null —
measured, the threshold was missed, and the entry is closed rather than
left open forever. A null that stays filed as pending debt is a claim
quietly waiting to be re-argued.
2. We publish honest nulls
Section titled “2. We publish honest nulls”Benchmark results — including the runs where the package changed nothing —
live in docs/benchmark.md. We do not delete a measured null to
make a number look better; the null is the evidence of honesty.
The freshest example: the persona-placebo benchmark
(2026-07-12, 3 arms × 2 providers, blind rubric judge, pre-registered
hypotheses) measured whether famous-figure persona framing improves decision
answers over the bare method text it wraps. It does not (Δ=0.17, p=0.607) —
while provider diversity moved judged quality ~15× more than persona identity.
The measured null closed a planned feature (persona panel-mode) instead of
shipping theater; the full per-cell data is committed at
internal/bench/reports/persona-placebo.json.
Behavioural-eval coverage — the honest baseline. Skill quality is only
as good as its measurement. Today 42 of 299 skills carry a behavioural
evals.json; the highest-traffic / highest-cost tiers (default-surface +
rich + routers) are fully covered (35 of 35), the long tail
(7 of 264) is not. We publish that gap rather than imply
“299 evaluated skills”: coverage is measured per tier
(./scripts-run src/scripts/skill_eval_coverage), CI-ratcheted so it can
only rise, and the priority tiers carry a hard tier floor: every
rich / default-surface / router skill MUST have a behavioural eval or an
explicit exemption-with-reason in internal/evals/tier-floor-exemptions.json
— no silent exemptions (the bright line deliberately replaces a
weighted-coverage score). Authoring long-tail evals is gated on per-case
human ratification (a generated assertion that checks the wrong property is
worse than none), so the number grows deliberately, not overnight.
Non-coding domain soundness — scoped, not proven. The finance /
founder / ops / content profiles sell concrete domain value (DCF,
runway, RICE, incident command, messaging), but the skills are forged on
TS/PHP — “promising, not proven” off those stacks. A disclaimer floor
bounds liability, not correctness: a skill can be format-correct,
disclaimered, and still embed a wrong domain assumption. Today 9 of 20
default-surface domain skills carry a sourced domain-truth fixture
(./scripts-run src/scripts/domain_soundness_status); the rest are labeled
unvalidated and the validated count is CI-ratcheted. The fixtures landed
so far are the deterministic targets (runway, unit-economics, DCF,
forecasting, scenario band/sensitivity), whose answer keys are computed
from cited standard formulas — never the skill’s own output; the
rubric targets (incident command, messaging, fundraising, editorial)
need domain-competent grounding and remain unvalidated, so validation
lands deliberately — the gap is published, never implied away.
Second-brain substrate — measured recall lift, honestly bounded. On a
deterministic multi-session recall corpus, the memory substrate beats a
no-memory baseline AND an equal-byte placebo: memory-on 27/27 vs
no-memory 10/27 vs placebo 9/27 (claude-haiku-4-5, 9 tasks × 3
seeds, sign test p = 0.031 for both pairings) — a real,
placebo-controlled lift (internal/bench/reports/second-brain-delta.json).
Scoped, not oversold: this is the context-value upper bound (perfect
retrieval on a one-fact-per-task corpus), NOT retrieval precision under a
large store; the lift concentrates exactly where memory is the only source
and ties where the prompt self-contains the fact. Boundary vs a human PKM,
and why the Obsidian export stays declined, are in
docs/second-brain-scope.md.
A follow-up removed the perfect-retrieval assumption: against a store of keyword-overlapping confusers the REAL retrieval recalls the needed decision into the top-5 (9/9) and the model disambiguates it — retrieval-on 27/27 vs no-memory 5/27 and vs placebo 5/27 (p=0.008 both). The honest limit: the keyword scorer recalls but does not rank (mean tie-set 3.3), which is the SQLite-FTS5 activation signal (ADR-116) at scale.
3. Known limits (published, witness-tested)
Section titled “3. Known limits (published, witness-tested)”Each limitation below is stated by the skill itself and carries a witness test that reproduces it. If the limitation is ever fixed, its witness goes red — so a “Known limit” can never quietly become false.
| skill | known limit | witness |
|---|---|---|
check-refs |
Only validates references to known-root paths (docs/, skills/, rules/, commands/, contexts/, personas/, …). A relative-path link such as ./sibling.md or ../foo.md is not matched, so a broken relative link is never reported. |
tests/scripts/witness/check_refs_relative_gap.test.ts |
4. What is checkable — us vs. the category
Section titled “4. What is checkable — us vs. the category”This is not a takedown — the point is the last column. For each claim,
our evidence is a pointer you can resolve on a fresh checkout; the wider
category is described only by what is publicly observable, never a named
competitor and never a counter-claim to anyone’s headline number. A claim
is “checkable” only when its our evidence pointer resolves — CI enforces
that (task check-comparison), so this column can never lie.
| Claim | Our evidence | The category | Checkable? |
|---|---|---|---|
| Capability/benchmark results are published — including the runs where the package changed nothing. | docs/benchmark.md |
A category headline figure (e.g. an ‘84.8%’ score) appears across marketing surfaces with no reproducible methodology published — so it cannot be verified either way. | ✅ |
| Every public claim binds to machine-checked evidence. | docs/CLAIMS.md |
Marketing claims in this category are not bound to a machine-checked ledger a reader can reproduce. | ✅ |
| Skills publish their own known limits, each with a witness test. | tests/scripts/witness/check_refs_relative_gap.test.ts |
Published known-limitation surfaces backed by reproducing tests are not a standard artifact in this category. | ✅ |
| A 30-second single-file wedge exists — one read-only subagent, one command, verdict with file:line — and its promise is a scoped, published eval, not a feature list. | internal/bench/orchestration/pv-a3-results.md#token_delta_provenance: measured |
Category entry points typically install the full platform (profiles, packs, runtime) before the first felt win; single-artifact entries with a published eval behind the promise are not a standard artifact. | ✅ |
| The persona-identity question was tested and published as an honest null (identity swap ~zero effect, provider choice ~15x larger) instead of being shipped as theater. | internal/bench/reports/persona-placebo.json#honest-null |
Named-persona prompts are shipped as a feature across the category without a published A/B separating identity from method. | ✅ |
| Every upstream-tool install prescription the package ships is version-pinned, intake-recorded, and enforced offline in CI — an unpinned version, a branch/archive install source, or a pipe-remote-to-shell instruction fails the build. | src/scripts/validate_reach_prescriptions.ts |
Capability layers in this category commonly install themselves and their upstream dependencies from a moving target (a branch archive or an unpinned package) via a piped remote installer; whether any given one pins is observable per repo, but a CI gate that fails the build on an unpinned prescription is not a standard artifact. | ✅ |
| Platform reads that the host’s own web tools cannot perform at all (a refused domain, an HTTP 402) are shipped as credential-free prescriptions whose reliability was measured per channel against a native control, with the thresholds frozen before the run and the narrowing caveat published alongside the pass. | docs/benchmark.md#ship-gated-reach |
Reaching gated platforms is commonly solved in this category by asking for an API key or a session cookie, and the capability is asserted from a feature list rather than a per-channel reliability measurement with a native-arm control; publishing the case where the native tools already suffice, next to the pass, is not a standard artifact. | ✅ |
| A code-intelligence engine was built natively, measured against plain disciplined grep, and then permanently defaulted off when it lost — the retrieval null (recall 0.365 vs grep 0.797 on graph-shaped questions) is published, and a deprecation date is on it. | internal/bench/reports/code-graph-vs-grep.md#Verdict — NULL, decisively |
Code-graph / semantic-index retrieval is shipped as a headline capability across this category; a published measurement against a plain-grep control — let alone one that retires the feature the vendor built — is not a standard artifact. | ✅ |
What each one prevents
Section titled “What each one prevents”The same rows, read as failure modes rather than as comparisons. Each names something that goes wrong without the control and carries the identical resolvable pointer — a projection of the table above, not a second list to keep in sync.
| Without it | The control | Evidence |
|---|---|---|
| You adopt on a headline number nobody can re-run, and never learn which of its measurements changed nothing. | Capability/benchmark results are published — including the runs where the package changed nothing. | docs/benchmark.md |
| A shipped number drifts from the artefact it came from and stays published, because nothing re-checks the claim against its source. | Every public claim binds to machine-checked evidence. | docs/CLAIMS.md |
| A limitation is discovered by you, in your repo, at the moment it costs you something — instead of being stated up front and held by a test. | Skills publish their own known limits, each with a witness test. | tests/scripts/witness/check_refs_relative_gap.test.ts |
| Evaluating the tool means installing the whole platform first, so the first thing you learn is the setup cost rather than whether it works. | A 30-second single-file wedge exists — one read-only subagent, one command, verdict with file:line — and its promise is a scoped, published eval, not a feature list. | internal/bench/orchestration/pv-a3-results.md#token_delta_provenance: measured |
| You pay tokens for prompt theater that was never measured against the method text it wraps. | The persona-identity question was tested and published as an honest null (identity swap ~zero effect, provider choice ~15x larger) instead of being shipped as theater. | internal/bench/reports/persona-placebo.json#honest-null |
| An install instruction silently follows a moving target — an unpinned package or a piped remote script — so what lands in your repo is whatever the source held that day. | Every upstream-tool install prescription the package ships is version-pinned, intake-recorded, and enforced offline in CI — an unpinned version, a branch/archive install source, or a pipe-remote-to-shell instruction fails the build. | src/scripts/validate_reach_prescriptions.ts |
| You hand over a credential — or your account — for a read the tools could already do, and you cannot tell which channels actually work because the capability was claimed rather than measured. | Platform reads that the host’s own web tools cannot perform at all (a refused domain, an HTTP 402) are shipped as credential-free prescriptions whose reliability was measured per channel against a native control, with the thresholds frozen before the run and the narrowing caveat published alongside the pass. | docs/benchmark.md#ship-gated-reach |
| You adopt an indexing engine on the promise that structure beats search, and nobody ever measured it against the search it replaces — so the index keeps costing build time and staleness risk for retrieval that got worse. | A code-intelligence engine was built natively, measured against plain disciplined grep, and then permanently defaulted off when it lost — the retrieval null (recall 0.365 vs grep 0.797 on graph-shaped questions) is published, and a deprecation date is on it. | internal/bench/reports/code-graph-vs-grep.md#Verdict — NULL, decisively |
4b. The two existing axes — enforcement level per rule, evidence form per claim
Section titled “4b. The two existing axes — enforcement level per rule, evidence form per claim”Pure projection of what the repo already knows — the enforced_by
resolution (check_enforcement_coverage) and the claims ledger
(docs/CLAIMS.md). No new taxonomy, zero hand-written rows.
Axis 1 — enforcement level per rule. 122 rules · 16 blocking (13.1%) · 10 observer · 0 local-only · 80 undeclared — of which 9 kernel-denied (block_kernel_rule_writes refuses an enforced_by write on a kernel rule, so no declaration is reachable for them at all) and 71 not declared yet.
denominator: 122 rule(s), frame in-scope (src/rules/*.md) == governed-total 122
| Rule | Effective level | Declared backstop(s) |
|---|---|---|
active-remediation |
none | instruction-only: the note, the ask and the user decision are all prose, so no gate can tell a discharged issue from a mentioned one |
code-comment-discipline |
validator | validator:src/scripts/lint_code_comments.tshook:comment-discipline |
code-provenance |
none | instruction-only: close-the-source-and-re-derive is a pre-write reasoning step only the model observes; CI checks the ledger, never the derivation |
context-hygiene |
observer | hook:context-hygiene |
council-availability |
none | instruction-only: no gate reads a chat claim about availability; check_council_config_location covers the tree side only |
decision-revisit-gate |
none | instruction-only: no gate can observe an agent citing a decision it never opened; adr_cite_check is deterministic where it runs and nothing makes it run |
design-fidelity |
none | instruction-only: no artifact in this tree records a fidelity comparison. lint_design_slop and lint_design_quality measure generic AI-aesthetic tells and accessibility; neither reads the handover, so a 1:1 claim is model-carried |
design-review-after-ui-write |
none | instruction-only: no artefact proves a design review happened outside the work-engine dispatcher; the review verdict is self-report |
evaluator-independence |
observer | hook:evidence-independence |
fix-what-you-see |
none | instruction-only: ownership-as-excuse is a disposition in prose; no gate can see a red check handed back with its cause named |
framework-neutrality-in-generic-skills |
validator | validator:src/scripts/lint_framework_leakage.ts |
git-history-discipline |
hook | hook:block-no-verify |
language-and-tone |
validator | validator:src/scripts/check_md_language.ts |
lethal-trifecta-guard |
validator | validator:src/scripts/lint_skill_frontmatter_safety.ts |
media-governance-routing |
validator | validator:src/scripts/lint_media_policy_linkage.ts |
minimal-safe-diff |
observer | hook:minimal-safe-diff |
missing-skill-recovery |
none | instruction-only: nothing can observe an agent concluding that no skill exists; the skill-route concern covers only the prompts where the ranker is confident |
neighbour-precedence |
none | instruction-only: the order is prose and a neighbour's prose can outvote it; only the listed effects hold |
no-roadmap-references |
validator | validator:src/scripts/check_no_roadmap_refs.tsvalidator:src/scripts/check_council_references.ts |
non-destructive-by-default |
none | none |
onboarding-gate |
observer | hook:onboarding-gate |
output-discipline |
validator | validator:src/scripts/lint_output_slop.ts |
persona-governance |
validator | validator:src/scripts/lint_persona_governance.ts |
playbook-precedence |
none | instruction-only: no gate can tell a playbook-first run from a skill-first one — both produce a diff, and which answer was consulted leaves no artefact |
preservation-guard |
validator | validator:src/scripts/check_condensation.tsvalidator:src/scripts/skill_linter.ts |
recurring-criticism |
none | instruction-only: the earlier disposition, the three outcomes and the hardening are all prose; the self-repair occurrence counter covers only detector-matched defects |
roadmap-progress-sync |
observer | hook:roadmap-progress |
secret-vcs-guard |
validator | validator:src/scripts/check_secret_leak.ts |
security-sensitive-stop |
none | instruction-only: threat-model-before-you-edit is a pre-edit reasoning step only the model observes |
self-repair-loop |
observer | hook:self-repair |
session-canary |
observer | hook:session-canary |
settings-ask-protocol |
none | instruction-only: no gate counts the questions in a chat turn; settings:set fences the illegal WRITE, never the badly-shaped ask |
skill-quality |
validator | validator:src/scripts/skill_linter.ts |
source-confidentiality |
validator | validator:src/scripts/check_no_external_sources.ts |
source-of-truth |
validator | validator:src/scripts/check_condensation.ts |
telegraph-speak |
observer | observer:maintainer-review |
test-first |
none | instruction-only: no gate sees which was written first; check_test_delta and check_test_weakening catch an absence, never an order. |
token-optimizer-maintenance |
validator | validator:src/scripts/check_token_optimizer_freshness.ts |
tool-safety |
validator | validator:src/scripts/lint_agent_security.ts |
ui-audit-gate |
observer | hook:design-pass |
untrusted-input-defense |
none | instruction-only: no deterministic gate inspects fetched content for injected instructions; injection_scan_hook is warn-only and default-OFF |
verify-before-complete |
observer | hook:verify-before-complete |
Undeclared rules (80, of which 9 kernel-denied) carry no row — an honest gap beats a false claim.
Axis 2 — evidence form per public claim. 104 ledger entries · 61 backed · 36 unbacked inventory · 7 resolved-null.
| Claim id | Kind | Status | Measured on | Evidence pointer |
|---|---|---|---|---|
adapter-lifecycle-day-one-table |
qual | backed | — | docs/contracts/provider-lifecycle.md#Current tier assignment |
adr-beneficiary-grade-bias |
quant | unbacked | — | PRE-REGISTERED 2026-08-21 (road-to-evidence-based-adr-governance Phase 6.2). WHY THIS EXISTS: it is the measured form of the conflict of interest a two-seat council named as the design's central risk. An agent grading a record that constrains it has a structural incentive to grade it weak, and recording provenance documents the conflict without neutralising it — an audit log is not a control. METRIC: partition accepted records into CONSTRAINING (the record forbids, gates, or requires human involvement in an action an agent could otherwise take) and NON-CONSTRAINING, classified from the record's Decision section by a party that did not assign its grade; compare the E0/E1 rate across the two partitions. THRESHOLD, fixed now: the constraining partition's weak-grade rate may exceed the non-constraining partition's by no more than 15 percentage points. POWER: at least 25 records per partition, else UNDERPOWERED. FALSIFICATION: (1) a breach is evidence of grade deflation and suspends any authority consequence outright — it does not trigger a re-grade by the same party, which would be the same conflict applied to its own audit; (2) the reverse skew (constraining records graded STRONGER) is reported as-is and is not read as a virtue: it would suggest defensive over-grading, which inflates reopen burden instead of deflating it; (3) a null result here does NOT license the authority coupling on its own — it removes one objection out of four (claim:adr-grade-accuracy-vs-gold, claim:adr-evidence-discovery-recallandclaim:adr-interruption-baseline are the others), and the coupling stays owner-reserved regardless. |
adr-evidence-discovery-recall |
quant | unbacked | — | PRE-REGISTERED 2026-08-21 (road-to-evidence-based-adr-governance Phase 6.2). WHAT IT MEASURES: the failure mode the discoveryfield exists for. A bare E0 collapses five states — evidence absent, evidence existed and was never cited, cited somewhere non-standard, present in the tree and not found, external and never fetched — and the last four are discovery failures, not evidence failures. A record graded weak because nobody looked is the cheapest possible way to manufacture a reopenable lock. METRIC: draw a random sample of at least 15 records carryingstrength: E0withdiscovery: complete, run a deeper independent search on each (full-tree grep for the decision's own terms, the roadmap and PR that produced it, the external sources its body names), and count how many turn out to have locatable evidence. THRESHOLD, fixed now: no more than 20% of the sample may turn out to have findable evidence. FALSIFICATION: (1) a breach means completeis being asserted whereincompleteis the honest value, and the remedy is to makeincompletethe only permitted value for a heuristic proposal rather than to lower the floor; (2) if nearly every honest answer isincomplete, that is itself the finding — the field then records uncertainty rather than resolving it, which is stated in ADR-240 § Assumptions as an accepted possibility rather than discovered later. HONEST-NULL PATH: if no record ever carries discovery: complete, the recorded finding is that the strong form was never used, and this claim resolves null with the observation that E0 in this tree means "unsearched" in practice. |
adr-grade-accuracy-vs-gold |
quant | unbacked | — | PRE-REGISTERED 2026-08-21 (road-to-evidence-based-adr-governance Phase 6.2). WHY NOT INTER-REVIEWER AGREEMENT, which is the obvious metric and the wrong one: agreement measures shared bias as readily as accuracy, and this repository has the receipt — 44 engine-shaped REJECT records accumulated under correlated council agreement (engine-reclassification-2026-07.md) and were disposed of by one measurement (claim:code-graph-retrieval-null). Two reviewers who searched the same way and read the same rubric will agree while both being wrong. METRIC: an anchor sample of 12-15 records is graded independently, then adjudicated to a gold value by a party that did not produce either grade; accuracy is the proportion of proposals matching gold on the E0/E1-versus-E2+ boundary, which is the boundary that matters because it is the one the burden table prices. THRESHOLD, fixed now: 85%. Reported WITH the disagreement count and stratified by record type, never as a bare percentage. POWER: fewer than 12 adjudicated records is UNDERPOWERED and yields no claim. FALSIFICATION: (1) high accuracy on records that constrain nothing, paired with low accuracy on records that constrain agent behavior, is a FAILURE even if the aggregate clears 85% — see claim:adr-beneficiary-grade-bias; (2) an adjudicator who saw a proposal first is not independent and that sample is void; (3) if adjudication itself proves unrepeatable, the finding is that evidence grading is not reliably gradeable here, which is a publishable null and closes the authority question by itself. |
adr-interruption-baseline |
quant | unbacked | — | PRE-REGISTERED 2026-08-21 (road-to-evidence-based-adr-governance Phase 6.1). REGISTERED BEFORE THE MECHANISM CHANGES ANY BEHAVIOR, and that ordering is the point: the axes ship descriptive (ADR-240 § 2 — a grade prices review burden and confers no authority), so the baseline is measured against a tree where nothing has been unlocked yet. METRIC, fixed now: a contact counts when a run stops or asks the owner AND the stated reason names an ADR — read from session transcripts, not from an agent's self-report, because "I was blocked by ADR-X" is exactly the claim the surrounding roadmap exists to stop taking on faith. DENOMINATOR: 20 roadmap runs, counted as /roadmap:process-* invocations that reach at least one closed step; a run that aborts before its first step is not a run. POWER: at least 20 runs on each side, else UNDERPOWERED and no claim in either direction — a five-run sample producing a favourable ratio is the failure this line exists against. FALSIFICATION: (1) a fall in contacts accompanied by a rise in the held-defect rate is NOT a win and is reported as a trade, not a success; (2) the comparison is rate-vs-rate, never absolute counts, since run volume will differ; (3) a rise is published as a rise — this metric can close the mechanism, and per ADR-240's own review_trigger a flat result reopens that record. HONEST-NULL PATH: if 20 runs do not accumulate, the recorded finding is that the metric was pre-registered and never had a population, and the interruption claim is withdrawn rather than left pending. |
adversarial-council-finding-coverage |
quant | resolved-null | — | docs/benchmark.md#adversarial-verification-council |
adversarial-review-tree-of-thoughts |
qual | backed | — | https://arxiv.org/abs/2305.10601 (2026-08-11) |
augment-manifest-version-package-synced |
qual | backed | — | src/scripts/lint_marketplace.ts#check_augment_manifests |
autonomous-analysis-self-refine |
qual | backed | — | https://arxiv.org/abs/2303.17651 (2026-08-11) |
budget-routing-relation |
qual | resolved-null | — | docs/contracts/budget-routing.md#Why it was retired |
bug-analyzer-chain-of-verification |
qual | backed | — | https://arxiv.org/abs/2309.11495 (2026-08-11) |
bus-factor-tracked |
qual | backed | — | docs/succession.md#trailing 90 days |
catalogue-pressure-null |
quant | unbacked | — | PRE-REGISTERED 2026-08-25 (road-to-routing-assurance Phase 3.3), quoted verbatim from the roadmap rather than paraphrased. DESIGN: the Phase 2 corpus run at N in {12, 20, 50, full}, distractors sampled deterministically in FNV-1a order -- the discipline rule_trigger_eval.ts:20-21and:147already use -- so the same seed reproduces the same distractor set. SCOPE, and it is a restriction the roadmap imposes on itself: this null "settles exactly one question, the confusion measurement, and cancels nothing". It may NOT be read as authority over tiering, because tiering already shipped for a different reason -- the host listing budget, viacompute_skill_tiers.ts. ON HOLD: if the null holds, record it and stop; no follow-up work item is created. ON BREAK: the result feeds the archived MCP roadmap's routable-skills-per-standing-token measurement rather than duplicating it. FALSIFICATION: a full-catalogue accuracy more than the floor delta below the N=20 figure, on a run whose distractor seed is recorded. |
code-graph-retrieval-null |
quant | backed | a build predating the 2026-08-22 extractor repair — post-repair recall UNMEASURED | internal/bench/reports/code-graph-vs-grep.md#Verdict — NULL, decisively |
command-count |
quant | backed | — | exec:check_artefact_count_messaging -> 0 |
conformance-advisory-vs-blocking |
quant | backed | — | src/domains/analysis-workbench/analyze/conformance/command.md#Both blocking carriers reached zero |
context-fidelity-compaction-compliance |
quant | unbacked | — | PRE-REGISTERED 2026-08-17 (road-to-context-fidelity Phase 0 Step 4 — registered while the census that would answer it, cf01, has NOT been run, so this is a genuine pre-registration and not a ledger entry written around a number already in hand). THRESHOLD fixed before data, taken verbatim from the roadmap: **a baseline compliance at or above 90 % for all three probe classes closes Phase 1 UNBUILT** and the null is published with the host version recorded. The three probe classes are a session-canary-bound obligation, a completion-gate reminder, and one trigger-loaded rule with a detectable obligation; the threshold binds all three separately, so a mean of 90 % carried by one strong class is not a pass. METHOD CONSTRAINT discovered before the census ran, and it is why this entry does not simply wait: cf03 measured 29 compaction events across 473 sessions, **all 29 tagged host-automatic and none tagged manual** (agents/evidence/eval-findings/context-fidelity-cf03.md). That zero is **absence of a RECORD, not absence of an event** — corrected on R2 finding 6, which caught the first phrasing here reporting an unobservable as an observation. The detector is pinned to one OBSERVED auto event (src/scripts/_lib/session_eol.ts:11-19) and nothing in the tree establishes that a manual compaction writes a compact_boundary record at all. So the constraint on cf01 is sharper than "measures a rare path": until manual detectability is established, a cf01 null is UNINTERPRETABLE — indistinguishable from a compaction that happened and left no trace. Establishing it is one manual compaction in one instrumented session, and it is a precondition rather than a result. FALSIFICATION fixed before data: (1) the host version is stamped on every observation, because compaction survival is a host fact that changes without notice; (2) a probe present only as paraphrase counts as NOT followed — the obligation is the behavior, not the recall; (3) at or above 90 % on all three classes the recorded consequence is that Phase 1 is not built and the folklore is named as folklore, with the same force as the positive direction. HONEST-NULL PATH is therefore the DEFAULT outcome of a high reading, not a fallback: this claim exists to be able to close work rather than to justify it. |
context-fidelity-memory-staleness |
quant | unbacked | — | PRE-REGISTERED 2026-08-17 (road-to-context-fidelity Phase 0 Step 4). THE THRESHOLD PREDATES THE DATA, THE LEDGER ENTRY DOES NOT, and the distinction is stated rather than blurred: the roadmap fixed **a stale ratio below 10 % shrinks Phase 2 to stamps only** on the day it was written, BEFORE any census ran; this entry was written after cf02 produced a first reading, so it records a resolution-in-progress rather than claiming the ledger row itself came first. FIRST READING, from agents/evidence/eval-findings/context-fidelity-cf02.md: 107 curated entries walked one by one against the tree — 73 still-true, 23 stale, 11 unverifiable, i.e. **21.5 % of all entries and 24.0 % of the verifiable subset**. Both denominators clear 10 %, so the kill criterion does not fire on either reading and the ladder stays justified. THE LOAD-BEARING FINDING is that the shipped instrument disagrees: memory_reportreportsstaleness-rate=0.0%because all 107 entries carry the SAMElast_validated: 2026-07-09and the SAMEreview_after_days: 365 — one bulk stamping event, so the age axis cannot read stale before 2027-07-09. Reading the kill criterion off that 0.0 % would have closed Phase 2 on a number that measures stamping rather than truth, which is the already-satisfied-test failure this repository has recorded before. WHY THIS STAYS UNBACKED despite having a number: the tree axis was walked BY HAND because no store-wide contradiction sweep exists (check_memory_contradictiontakes–type –key –body, i.e. it validates one proposed entry), three observers each classified one store, and inter-rater agreement is therefore UNMEASURED. A hand classification is not a reproducible instrument, and a ratio that cannot be re-derived by a command is not backing. BACKING TRIGGER: a store-wide sweep exists AND reproduces a ratio within its own stated error of 21.5 %. FALSIFICATION fixed before data: (1) unverifiable entries are counted as their own class and folded into neither side — folding them into still-true inflates the pass rate, folding them into stale manufactures defects; (2) the commit anchor is the precondition for any automated reading, because without it a date cannot be tied to a tree state; (3) below 10 % on a reproducible sweep the recorded consequence is the roadmap's own: the ladder is unbuilt, only stamps ship, and the null is published. |
context-token-reduction |
quant | backed | — | internal/bench/reports/token-baseline.json#eager_rule_load |
council-fallback-loses-zero-seats |
qual | backed | — | exec:vitest run tests/scripts/ai_council/council_cli.test.ts -> 0 |
council-parse-outcome-corpus-rate |
quant | backed | — | exec:vitest run tests/scripts/ai_council/parse_corpus.test.ts -> 0 |
council-vs-solo-baseline |
comparative | unbacked | — | PRE-REGISTERED 2026-07-12 (road-to-feedback-8.11 Phase 3 — no goalpost-moving after the numbers land; design at docs/design/council-vs-solo-baseline.md). Falsification criteria fixed BEFORE data: (1) quality = blind post-hoc grading against known ground-truth dispositions, two blind judges, admissible only at Cohen's κ ≥ 0.60 (reuse check_quality_regression.ts kappa machinery); (2) the five feedback-proposed admission dimensions are recorded per decision AT pre-registration, so "≥2-of-5" is a testable post-hoc correlate, never a pre-imposed gate; (3) NO lift on any subset (overall, per impact class, per dimension stratum) → honest null, deliberation-protocol phases stop (maintenance-only), recorded in road-to-opt-council-deliberation; lift on a subset → admission criteria derived FROM that subset's characteristics. Execution is spend-gated (user confirms rendered estimate in-session); shadow-log was absent/empty at design time — zero prior council-vs-solo data exists. |
critic-protocol-load-bearing-ab |
quant | resolved-null | — | internal/bench/adversarial-council/runs/critic-protocol-ab-report.json |
cross-model-parity-count |
quant | backed | — | internal/bench/reports/parity-count.json#min over hosts of median |
cross-source-consistency-precision |
quant | unbacked | — | PRE-REGISTERED 2026-07-28 (road-to-feedback-9.2.0-followups Phase 1 — no goalpost-moving after the numbers land). Falsification criteria fixed BEFORE data: (1) fixtures + expected actions are pinned in internal/bench/corpora/honesty-false-premise.yaml(shared with the honesty bench, extended-not-forked); (2) the scorer issrc/scripts/bench_cross_source_eval.ts (ask|proceed|warn classification, forbidden-assumption + over-firing checks) — precision = correctly-surfaced discrepancies / all surfaced; over-firing = asks on negative controls / negative controls; (3) the run needs real model responses per fixture (paid, maintainer-gated spend) — this entry stands as documented debt until that run lands; (4) HONEST NULL consequence bound: precision < 85% or over-firing > 5% → loosen the rule's default (consistency.cross_source: on→auto) or tighten its confidence tiers — never silently keep firing. This binds the weaker-evidenced default-on rule to a measurement like every other default-flip. |
default-install-context-cost |
quant | backed | — | exec:update_counts --check -> 0 |
delivery-path-parity |
quant | unbacked | — | PRE-REGISTERED 2026-08-25 (road-to-routing-assurance Phase 4.2), quoted verbatim. EPSILON: 0.05 absolute recall, fixed in docs/contracts/routing-assurance-metrics.md before any parity run. DESIGN: one corpus file, identical prompts and floors, parametrized over host-native listing and MCP-tool listing -- both paths exist in the tree today, so this has no external dependency. CONSEQUENCE OF A BREACH, pre-registered so it cannot be softened afterwards: a breach "blocks any MCP default-on decision; default-off holds until then". PUBLICATION: the parity table is published as evidence against the archived MCP roadmap's measured-null outcome and adds NO new claim id -- which is why this entry covers the parity gate and not the table. FALSIFICATION: a measured delta worse than -0.05 on the full Phase 2 corpus, or a corpus that cannot be run on both paths, in which case the claim is UNDERPOWERED rather than broken. |
description-gate-catches-regressions |
quant | unbacked | — | PRE-REGISTERED 2026-08-25 (road-to-routing-assurance Phase 0.5), BEFORE description_route_checkexists. BASELINE: zero by construction -- no gate reads the description surface today. The deterministic suites testdist/router.json trigger substrings (trigger_coverage.ts:10) and the 94 routing-matrix fixtures; production skill selection runs on SKILL.md description (lint_skill_descriptions.ts:6-7, "the agent picks a skill from its description"); and the only harness on that surface is rule_trigger_eval.ts, which is "advisory only, never gating" (:4) and whose "live floor breach fails the SCHEDULED canary job only -- PRs are never blocked by live results" (:32-33). METRIC: on a corpus of description edits with known direction, the fraction of regressions the gate blocks (recall) and the fraction of neutral edits it does not block (precision). THRESHOLD: pre-registered per unit in docs/contracts/routing-assurance-metrics.md as the Phase 0.2 baseline minus a fixed 0.10 absolute tolerance, with the tolerance FIXED BEFORE the baseline run so it cannot be tuned to a result. RECALL-FIRST: per Phase 1.3 a positive that stops loading is the failure that matters, so the fail condition is recall-shaped and a precision miss is reported rather than blocking. FALSIFICATION: the gate blocks no seeded regression, or it blocks neutral edits at a rate that makes it unusable at PR time. STATED LIMITATION, not discovered later: the checker is a PROXY -- it asks whether a description is distinguishable from its neighbours, not whether a production model selects it. A green gate is evidence that a description did not get LESS distinguishable, never that production routing works. |
design-slop-false-positive-baseline |
quant | backed | — | internal/bench/corpora/design-slop-fp-PREREG.md#The ceiling, declared before the run |
discipline-lift-weak-host |
quant | backed | — | docs/benchmark.md#weak-host-specific |
dispatch-event-capture-reliability |
quant | resolved-null | — | PRE-REGISTERED 2026-08-30 (road-to-experience-loop-broadening`` |
domain-soundness-scoped |
qual | backed | — | exec:domain_soundness_status --check -> 0 |
domain-soundness-validated-count |
quant | backed | — | exec:domain_soundness_status -> 0 |
downshift-cost-reduction |
quant | backed | — | internal/bench/routing-downshift/results-2026-07-08.md#FAMILY-SCOPED PROVE |
encoding-corpus-scope-guard |
qual | backed | — | tests/scripts/encoding_corpus.test.ts#FAILS when a deliberately out-of-scope fixture is added |
encoding-floor-text-layer-only |
quant | backed | — | agents/evidence/reports/encoding-floor-measurement.md#Selected branch: ADOPT |
enforcement-coverage-resolved |
quant | backed | — | exec:check_enforcement_coverage --check -> 0 |
enforcement-undeclared-denominator |
quant | backed | — | exec:check_enforcement_denominator -> 0 |
essential-tier-cost-factor |
quant | backed | — | docs/benchmark.md#REPLICATION FAILED |
eval-coverage-ratcheted |
qual | backed | — | exec:skill_eval_coverage --check -> 0 |
experience-loop-repeated-failure-effect |
quant | unbacked | — | PRE-REGISTERED 2026-08-30 |
experiment-loop-iteration-floor |
quant | unbacked | — | PRE-REGISTERED 2026-08-17 (road-to-metric-loop-and-review-integrity Phase 0/5 — the floor was fixed before the spike ran, and the spike's kill criterion was its complement: "fewer than five clean iterations" would have left Phase 3 unbuilt). Measured on a toy metric in a scratch repository, agents/evidence/eval-findings/metric-loop-s01.md: 6 clean iterations, metric 24 → 3, with iteration 5 reverting a change that improved the metric 67 % and broke behavior. That result answers the PHASE GATE and is why the skill shipped. It does NOT back this claim, and the distinction is the whole reason the entry stays unbacked: the run was one agent, one session, one toy metric whose evaluator was written alongside the loop, so it measured whether the PROTOCOL holds, not whether the shipped skill drives a real metric. BACKING REQUIRES: ≥ 3 runs of the shipped experiment-loop skill against metrics that existed before the run, each with its register committed, each reaching ≥ 5 clean iterations. DROP: any run below the floor publishes the null and the skill is withdrawn rather than the floor lowered — lowering a pre-registered floor after seeing the data is the tuning this roadmap's own s04 finding forbids. |
forensics-pack-value |
quant | unbacked | — | agents/evidence/release-findings/ |
free-jury-judge-agreement |
quant | unbacked | — | PRE-REGISTERED 2026-09-07 (road-to-admissible-council-seats Phase 3.1), BEFORE any measurement run and before any seat capable of running it is authorized -- the roadmap's free-seat-measurement-spendblocker was DECLINED for this round in the same change, so the claim is filed with no result and no route to one. That ordering is the pre-registration: a threshold fixed while the measurement is impossible cannot have been fitted to a result. THRESHOLD: Cohen's kappa >= 0.60, the conventional "substantial agreement" floor, computed with the existingcohensKappainsrc/scripts/check_quality_regression.ts:117rather than a new statistic written for this claim. CORPUS: repo-owned PUBLIC artifacts only -- skill descriptions, the trigger corpus, rule text, bench fixtures -- never consumer content; the seat-level enforcement is thepublic-artifactcontent ceiling insrc/scripts/ai_council/content_ceiling.ts. PANEL: at least two model families, enforced by juryAggregateinsrc/scripts/ai_council/jury_aggregate.ts, which returns absent for a single-family panel rather than scoring it -- docs/CLAIMS.mdalready records that a same-posture second vendor's catches were a strict SUBSET of the first's, so a same-family panel measures redundancy and not agreement. AGGREGATE: a trimmed mean or median, never a vote count, for the same reason. POSITION ORDER swaps per judge viasrc/scripts/ai_council/judge_position_bias.ts. ONE CLAIM, NOT TWO: the proposer-lift question is held behind metered-backend-parkinagents/roadmaps/later/road-to-governed-evidence-production.md and is deliberately NOT filed here. FALSIFICATION: a measured kappa below 0.60 on a run whose panel composition, corpus slice and swap seed are recorded. HONEST-NULL PATH, written before the run rather than after it: below the bar the jury stays evaluation-only and NO gate consumes its output; lifting that is not authorized by this claim and requires a new pre-registered claim, which is the scoped form -- a refusal that said "permanently" would prejudge a future owner ruling. ROLE BOUND, independent of the result: ADR-257 records that a route this package did not pay for may propose and may score but may never carry a verdict, chair, or act as evaluator of record; a passing kappa does not lift that boundary, because the boundary is about accountability and not capability. UNDERPOWERED is neither a pass nor a null: a corpus slice too small to separate 0.60 from chance settles nothing and may be cited for neither direction. |
gated-platform-reads |
quant | backed | — | docs/benchmark.md#ship-gated-reach |
governance-adjacent-properties |
quant | backed | — | internal/bench/reports/governance-invariants.json#adjacent_properties |
governance-aggregation-refusal-invariance |
quant | backed | — | internal/bench/reports/governance-invariants.json#s0_1_aggregation_steerability |
governance-decomposition-effect-boundary |
quant | backed | — | internal/bench/reports/governance-invariants.json#s0_2_decomposition_laundering |
governance-marker-preservation-null |
quant | backed | — | internal/bench/reports/governance-invariants.json#s0_3_marker_survival |
hook-dispatch-latency |
quant | backed | — | docs/hook-latency.json#invocation_path |
host-agent-count |
quant | backed | — | exec:vitest run tests/install/toolDetection.test.ts -> 0 |
humanizer-tell-reduction |
quant | backed | — | internal/bench/reports/humanizer-v1.md#prefers the humanized text |
injection-scan-corpus-rates |
quant | backed | — | agents/evidence/reports/injection-detector-wiring.md#The numbers |
install-audit-clean |
quant | backed | — | .github/workflows/release-validation.yml#npm audit --omit=dev --audit-level=high |
judge-family-llm-as-a-judge-foundation |
qual | backed | — | https://arxiv.org/abs/2306.05685 (2026-08-11) |
keyword-anchoring-census |
quant | backed | — | agents/evidence/analysis/anchoring-independent-replay-2026-08.md |
lean-init-cost-reduction |
quant | unbacked | — | PRE-REGISTERED 2026-07-28 (road-to-lean-agent-init Phase 3 — registered BEFORE any savings number is cited anywhere; family-scoped, modeled on downshift-cost-reduction; quality definition reused from the correctness-comparison acceptance, no second truth). Falsification criteria fixed BEFORE data: (1) correctness floor — primitive answer ≡ agent answer on the golden corpus (internal/bench/lean-init/results-2026-07-28.md, 12/12); ANY mismatch on a routed real task recorded via correctness_match: false counts against the claim; (2) negative control — a non-lookup task never routes to a primitive (LOOKUP_CORPUSlk-n1..n4, FP=0); (3) cost metric — read fromagents/runtime/state/audit/*.jsonlorchestration lines taggedorigin: lean-init-2026withlookup_class != null, comparing route_taken: primitivetoken cost againstroute_taken: subagentlines of the same class (n and family scope stated at backing time); (4) segregation — lines carryorigin: lean-init-2026so theroad-to-orchestration-scope-decision sample stays uncontaminated (council Q5, 2026-07-28). PROVE → flip to backed for the lookup family only; DROP → honest null, primitives stay (correctness-validated) but no savings number is ever cited. |
ledger-exec-verifiability |
quant | backed | — | internal/reports/exec-evidence-feasibility.json#"exec_feasible" |
lexical-ranking-lift |
quant | backed | — | exec:measure_lexical_ranking -> 0 |
mcp-registered-server-standing-cost |
quant | backed | — | agents/evidence/metrics/mcp-tool-standing-cost.jsonl#tool_search_threshold |
no-runtime-daemon |
qual | withdrawn | — | docs/contracts/no-runtime-boundary.md |
obligation-settle-shadow-bar |
quant | unbacked | — | PRE-REGISTERED 2026-09-13 (road-to-a-ledger-that-closes-the-loopsteps 5.2 and 5.3), committed in the SAME change that ships the detector and BEFORE any code able to refuse exists. The concern returnsEXIT_ALLOWon every path and the manifest declares itseverity: advisory, fail_closed: false, so at the moment this bar is filed there is no arming switch to fit it to. That ordering is the pre-registration. |
orchestration-dispatch-net-win |
comparative | unbacked | — | PRE-REGISTERED 2026-07-11 (road-to-orchestration-scope-decision Phase 1 — no goalpost-moving after the numbers land). Falsification criteria fixed BEFORE data: (1) held quality is deterministic, scored by src/scripts/check_quality_regression.tsthresholds — a token/wall win that degrades output below the regression threshold FAILS the claim; (2) negative control —pv-02-negative-controlmust NOT trigger dispatch (a classifier that fires on everything is a cost leak, not a win); (3) win metric — ≥15% reduction in token-or-wall onorch-02+orch-03vs the single-agent baseline, read fromagents/runtime/state/audit/*.jsonlorchestration lines throughgateVerdict()/resolveShippedDefault(). Binds to a resolving report once ≥20 real ask-mode telemetry lines exist (Phase 2 — maintainer-run; the corpus –runagent-spawn is gated out of auto-mode). PROVE → flip to backed for the proven family only; DROP → renewed honest-null, keepask, demote orchestration from the public value proposition. |
orchestration-observed-dispatch-cost |
comparative | resolved-null | — | internal/bench/orchestration/backfill-2026-08-07-verdict.md#honest null |
persona-identity-placebo-null |
quant | backed | — | internal/bench/reports/persona-placebo.json#honest-null |
plaintext-source-attribution |
qual | backed | — | exec:check_no_external_sources -> 0 |
plan-gates-measurement-protocol |
quant | unbacked | — | docs/contracts/plan-review-gates.md#Advisory window (Stage A, verdict #20) |
positioning-honest-nulls |
qual | backed | — | docs/benchmark.md#honest |
provenance-detector-transformation-sensitivity |
quant | resolved-null | — | PRE-REGISTERED 2026-07-28 BEFORE the S0.3 baseline run (road-to-provenance-and-license-governance Phase 0; corpus frozen at content-sha256 dbbc84a7325e4fa38483ba05d35d9c0c98fa822ae25d873bd5efbafaf2534bb3 over internal/bench/provenance/, 36 files). Thresholds fixed BEFORE data, per the roadmap's S0.2 and its denominator fix: (1) detector recall on the verbatim+rename-only subset >= 10/16 (8 verbatim + 8 rename-only); (2) false positives on the 12 independent controls <= 1/12; (3) rename-only samples MUST hit (principle 6 — laundering by rename cannot clear a hit); (4) structural-rewrite samples form the residual class and their recall feeds the Phase-5 drop gate (>= 21/24 on the full seeded corpus DROPS Phase 5). The floor is a GO/NO-GO gate for building the CI layer, never the marketed capability — the marketed capability is the measured rate published per S3.1/S3.3 with the scope bound above co-located. HONEST-NULL consequence (K1): thresholds missed => no deterministic-gate claim ever, the behavioral layer ships alone, null published; no silent threshold adjustment. |
provenance-gate-effectiveness |
qual | unbacked | — | PRE-REGISTERED 2026-07-28 (road-to-provenance-and-license-governance Phase 3, S3.1 — registered AFTER Gate G0's verdict, so this claim's text already reflects the re-scope rather than describing a capability that was later cancelled; the original S3.1 draft text ("AC's provenance gate detects seeded verbatim and rename-only OSS copies at the Phase-0 measured rate") is FALSE post-G0 and is not reused). G0 honest-null context (see provenance-detector-transformation-sensitivityfor the pre-registered thresholds): the deterministic scan layer (jscpd offline + SCANOSS online) measured, on the frozen synthetic corpus, verbatim+rename-only recall 12/16 (union) and false positives 2/12 (union) — missing BOTH the recall and FP thresholds — with SCANOSS alone recalling rename-only samples 0/8. Council decision 2026-07-28 (K1 literal, Option A): nolint_code_provenance.tsships in ANY form in CI, not even advisory — the scan capability exists ONLY as thelicense-compliance-auditskill (src/skills/license-compliance-audit/), invoked deliberately by a human, never wired into any pipeline. Falsification criteria fixed at registration: (1) any deny-class or unknown-license ledger entry passesci(alint_provenance.tsregression); (2) any ledger entry missing atransformation_notepassesci; (3) any rename-only-phrased transformation_note(the 15-phrase rejection list) passeslint_provenance.ts; (4) any user-facing surface asserts or implies a CI-facing similarity/duplication detector exists, or omits the co-located scope statement (S3.2/S3.3 global consequence bound, enforced by lint_provenance_vocabulary.ts). Backing trigger: (a) >=1 real (non-empty, non-fixture) ledger entry survives lint_provenancein a merged PR with a substantive transformation_note, AND (b) the Phase-4 dogfood self-audit (S4.1) publishes its findings against this repo itself, whatever the outcome. Part (b) is already satisfied —internal/bench/provenance/reports/self-audit-2026-07-28.mdpublished a headline finding that 551 of 552 online-scanner hits on this repo's OWN source were self-matches against its own published releases, a third independent argument for the G0 verdict and evidence the corpus's 2/12 false-positive rate understated the real-world surface. Part (a) is NOT yet satisfied —provenance/borrows.jsonl holds zero entries — so this claim stays unbacked until a real borrow lands and clears the ledger. |
published-artifact-counts |
quant | backed | — | exec:update_counts --check -> 0 |
red-before-production-edit-rate |
quant | unbacked | — | PRE-REGISTERED 2026-08-26 (road-to-evidence-gated-change Phase 6.1 -- no goalpost-moving after the numbers land). INSTRUMENT: agents/runtime/state/test-results.json, written by src/scripts/_lib/test_red_state.ts, which records the target, the observed failure CLASS and a run identifier. A run counts toward the numerator only when the record carries one of the three VALID red classes (assertion, missing-target, contract) for the target under change; a broken fixture, a test syntax error, a missing unrelated dependency and a runner fault are recorded and explicitly do NOT count. DENOMINATOR: work-engine runs that made at least one production edit. BASELINE: none exists -- the instrument was created by this change, so the pre-period is UNMEASURABLE and no before-number is claimed. The first reading is therefore a level, not a delta, and a second reading after a comparable period is what makes the direction claimable. FALSIFICATION: (1) the rate does not rise, which is an HONEST NULL and closes the claim as measured-null rather than reopening the instruction wording; (2) the rate rises while the record shows mostly missing-target for targets nobody later implemented, which would be the instrument being gamed rather than the discipline improving, and voids the reading; (3) the record proves per-machine and gitignored, so a cross-machine claim is out of scope by construction and is not made. WHY ONLY ONE CLAIM: this is the one metric an instrument exists for after Phase 4.1. The source proposals asked for several more -- duplication rate, reuse-verdict accuracy, time-to-green -- and none of them has an instrument in this tree, so they are named here as unmeasured rather than pre-registered. DATED 2026-10-06: the instrument's own registry row (src/config/assurance-capability-registry.json, test-red-evidence) was corrected from availabletodegradedthe same day --test_red_state.ts's writer has no production caller, so no run populates test-results.jsontoday and this claim's denominator is currently always zero. The claim staysunbacked, which was already the honest state; the instrument it cites is not yet the one it describes. |
reference-loop-upgrade-value |
quant | unbacked | — | PRE-REGISTERED 2026-08-12 (road-to-cross-repo-differential-loop Phase 6 — no goalpost-moving after the runs land). Falsification criteria fixed BEFORE data: (1) the comparison is against a shadow run of the pre-upgrade command text on the same reference, so "could not have produced" is decided by diffing two documents, not by assertion; (2) a finding counts only with a concrete file:line on OUR side — a probe recording consumer not locatableis an honest result but not a positive; (3) TIME BOUND: 180 days from merge — an event-bound measurement on a rare event is an unbacked row that never settles, so window expiry counts as the bar not cleared; (4) HONEST NULL consequence bound, asymmetric by construction: bar not cleared or window expired → the interop-probe, convergence and–deep mechanisms revert and the null is published; the anchor-table and bound-claim-gate mechanisms stay regardless, because they enforce ADR-211 C/D and the claims ledger — doctrine that already binds — rather than claiming new value. A measurement that never arrives therefore cannot leave the command worse than before the upgrade. |
release-hold-refuses-declared-state |
quant | unbacked | — | PRE-REGISTERED 2026-09-13 (road-to-release-holds-that-refuse Phase 0.4), BEFORE any evaluator, marker grammar or template rule exists -- Phase 1.1 is itself still blocked on an owner approval (rule-13-amendment), so nothing in the tree can yet carry the state this claim is about. That ordering is the pre-registration: a threshold fixed while the mechanism is unbuildable cannot have been fitted to a result. BASELINE, measured not assumed: zero, by construction. src/scripts/release.tsis 1,814 lines with exactly one occurrence of the wordroadmap (:322, inside a comment), and the roadmap template's rule 13 forbids a roadmap from saying anything about shipping at all -- so no release boundary reads a roadmap and no roadmap may address a release. Full Phase 0 sweep at agents/evidence/analysis/release-holds-phase-0-2026-09-13.md, taken at 7182f5d07. DENOMINATOR: 30 consecutive release tags cut after Phase 4 lands, counted from the first tag whose tree carries a wired refusal; tags are the denominator rather than calendar time because a refusal can only fire at a cut. NUMERATOR, either arm counts: (a) at least one release refused by the hold evaluator with the refusal logged, or (b) at least one roadmap that would have carried a hold re-sequenced to a continuous phase shape at authoring time, logged by the Phase 1.4 authoring self-check. Arm (b) exists because the mechanism's intended effect is that holds are RARE -- a hold that is never needed because authors re-sequence instead is the mechanism working, and counting only arm (a) would score that outcome as a failure. FALSIFICATION: 30 tags pass with zero refusals AND zero logged re-sequences, in which case the claim resolves to an honest null. EXPECTED OUTCOME, stated before the run rather than after it: the null is the PREDICTED result, not a surprise. The active corpus holds zero release-coupling declarations at HEAD, 0 of the 7 status: ready roadmaps is mid-flight, and the migration candidate the source proposal named is archived with both of its coupling risks closed (MITIGATED / DISCHARGED); at tag 15.0.0 the exposure population was 3 of 7 active roadmaps mid-flight -- and none of those three declared a coupling either, so even that population is one of exposure and never of violations. ON NULL: Phase 6.2's disposition holds -- the primitive stays and only the free parts (the boundary screen line, the runbook bullet) are struck, with the tag range recorded beside template rule 28 (renumbered 27 -> 28 on 2026-09-14 when an unrelated rule 27 landed; this reference was the one cross-reference the renumber missed). Deleting the primitive is NOT the null disposition; the owner's constraint is that a broken intermediate state must not ship by accident, and a mechanism whose value stayed latent is not one that failed. SCOPE: this measures whether a declared unpublishable state is refused at a cut. It says nothing about whether roadmaps declare such states HONESTLY -- the marker is only as true as the checkbox flip behind it, which is instruction-only and is recorded as risk rank 3 on the roadmap rather than claimed as solved. UNDERPOWERED is neither a pass nor a null: fewer than 30 post-Phase-4 tags settles nothing and may be cited for neither direction. |
resident-process-permitted-under-governance |
qual | unbacked | — | docs/decisions/ADR-249-supervised-resident-process-permitted-under-governance.md |
retrieval-substrate-live-pass |
quant | backed | — | internal/bench/reports/second-brain-retrieval.json#retrieval-on |
review-independence-changes-consumption |
qual | unbacked | — | PRE-REGISTERED 2026-08-17 (road-to-metric-loop-and-review-integrity Phase 2/5 — registered BEFORE any consumption claim is made anywhere). The MECHANISM shipped and is machine-checked: check_review_schemarefuses an artifact whoseacceptance_statuscontradicts itsreview_independence, and its –self-testplants the exact defect (a same-family set claimingaccepted) and confirms the rejection fires. The EFFECT is a different question and is not measured: nothing yet observes a consumer reading the field and behaving differently, and the only committed ledger (9.14.0) was backfilled by this same change rather than consumed by anyone. BACKING REQUIRES: ≥ 2 recorded instances where a reader or a downstream gate declined to treat a provisional artifact as acceptance, with the artifact and the decision both citeable. DROP: if the fields ship for one release and every consumer still reads the verdict line alone, the honest null is that the metadata is inert — the Risk-Register rank-2 outcome — and it is published as such rather than defended. |
roadmap-wall-clock-baseline |
quant | unbacked | — | PRE-REGISTERED 2026-08-17 (road-to-user-out-of-the-loop Phase 0 Step 3). SEPARATE FROM user-out-of-loop-baselineBY CONSTRUCTION, never merged into it: the roadmap's Goal states the two axes are deliberately not one, because a run can ask zero questions and still be slow — a single blended metric would let a contact win pay for a wall-clock loss and report the pair as progress. INSTRUMENT:src/scripts/interruption_report.tsderives elapsed time per run fromagents/runtime/.agent-chat-historytimestamps and splits it into WAITING (agent turn → the next real user turn) and WORKING (elapsed minus waiting). The join to the contact axis is the session tag: the ledger writesrun_idviaderive_session_tag, the same derivation that file writes as s. SYNTHETIC-TURN EXCLUSION is part of the definition, not a filter applied later: the harness writes task notifications and system reminders into the user role, and counting those as replies collapses every measured wait toward zero and makes the whole axis read as already-solved. The count of excluded turns is reported so the exclusion is auditable. POWER CAVEAT: identical to the sibling claim — 5 sessions measured against a 30-session request on the day of registration; window_shortis reported and a short window may not be cited as a baseline. FALSIFICATION fixed before data: (1) the QUALITY ANCHOR is the held defect rate, same as the sibling — a wall-clock win that moves it is a FAIL; (2) WORKING time, not elapsed, is the number a mechanism is judged on when the mechanism claims to remove waiting — reporting an elapsed improvement produced entirely by a faster human is the attribution error this criterion exists to block; (3) ≥ 20 recorded runs before any comparison, else UNDERPOWERED. HONEST-NULL PATH: if elapsed falls while working time does not, the recorded finding is that the change moved the human's response time and not the run's, and no autonomy claim is made from it. POST-REGISTRATION FINDING 2026-08-19 (road-to-long-horizon-execution, added AFTER registration and changing NO threshold — the ≥ 20 floor above stands exactly as written): the floor is **structurally unreachable with this instrument at default retention**, which is a different statement from "not yet reached" and has a different remedy. Timing comes only fromagents/runtime/.agent-chat-history, whose retention is DEFAULT_MAX_SESSIONS = 5 (src/scripts/chat_history.ts; chat_history.max_sessionsis unset on every settings layer, so the default is live). Five retained sessions yielded **4** timing-bearing runs on 2026-08-19, and the same file held 5 sessions at registration on 2026-08-17 — two readings two days apart, both at the cap. Waiting therefore does not fill this window; it rotates it. Backing this claim requires either a timing source that is not a rolling buffer, or an explicit retention change with its own privacy review — or this claim closes on the honest-null path above. Recorded here rather than in a roadmap because the reachability of a pre-registered floor is a property of the claim, and a reader deciding whether to wait for more data needs it at the point of the claim. The siblinguser-out-of-loop-baselineis NOT affected: its source is the committed append-only ledger, which stood at **19** of 20 on the same day and is reachable by one more recorded run.interruption_reportnow prints each axis's own N against the floor, because the single ⚠️ SHORT WINDOW banner over both axes had already produced one live misreading —runs: 21 read as the contact axis clearing the floor, when 2 of those runs carry timing and no ledger entry. |
rule-count |
quant | backed | — | exec:check_artefact_count_messaging -> 0 |
scope-dedup-cold-start-reduction |
quant | backed | — | agents/settings/contexts/cache-economy-refusals.md#Honest null — scope de-duplication is measured but |
scoped-dangle-follow-rate |
quant | unbacked | — | PRE-REGISTERED 2026-08-23 (road-to-skill-link-integrity-and-manifest-sync Phase 4). POPULATION, measured and reproducible: 24 dangling links from 17 surviving skills, derived by lint_handoffs –census-jsonwithis_pruned_under_scoped— the predicateinstall.tsitself applies, so the counted set cannot describe a projection the installer does not perform. METRIC: read attempts against.claude/skills/over a 30-day window, fromagents/runtime/metrics/skill-usage.jsonl. THRESHOLD, fixed now: zero attempts over a LIVE window closes this as a published null and the 24 links stay; a nonzero count promotes the fix, which is to rewrite each dangling link in the PROJECTED SKILL.md to name the slug and its pack instead of linking it — source tree untouched, using the same predicate the counter uses. INSTRUMENT STATUS: **dead, and the measurement was therefore not attempted.** Two independent reasons, both verified: (1) the store is gitignored and machine-local, so it is ABSENT in any fresh checkout, worktree, or CI run — that is the state the committed row agents/evidence/metrics/scoped-dangle-follow-rate.jsonrecords; in the maintainer's parent checkout it holds 181 records whose newest timestamp is 100 days old (2026-05-15T13:44:17.594Z), so it is stale there rather than absent. (2) **A LIVE clock would still not answer this**, which the drafting phase did not foresee: every one of those 181 records carrieskind: “exposure”, and no event in FOLLOW_KINDS (read, read_attempt, follow) is emitted anywhere in the tree — the instrument records that a skill was SHOWN, never that a link was FOLLOWED. BLOCKED ON: emitting a follow event, which is not in this roadmap. FALSIFICATION: attemptsisnulland never0wheneverinstrument_liveis false, asserted intests/scripts/scoped_dangle_window_guard.test.ts; a 0 there would be the false null this whole phase exists to prevent, and reporting one is the failure, not the finding. |
second-brain-recall-lift |
quant | backed | — | internal/bench/reports/second-brain-delta.json |
sequential-thinking-chain-of-thought |
qual | backed | — | https://arxiv.org/abs/2201.11903 (2026-08-11) |
shipped-artifacts-hidden-instruction-scanned |
qual | backed | — | exec:lint_agent_security -> 0 |
skill-activation-census-zero |
quant | backed | — | exec:check_skill_activation_claim -> 0 |
skill-count |
quant | backed | — | exec:check_artefact_count_messaging -> 0 |
skill-improvement-reflexion |
qual | backed | — | https://arxiv.org/abs/2303.11366 (2026-08-11) |
skill-link-census |
quant | backed | — | agents/evidence/metrics/skill-link-census.json#"dead_links": [] |
skill-tiering-h1-unmeasured |
quant | backed | — | agents/evidence/analysis/skill-tiering-matrix-arm.md#The question this arm cannot answer |
skill-tiering-h2-costs-more-by-default |
quant | backed | — | agents/evidence/analysis/skill-tiering-matrix-arm.md#H2 |
subagent-valid-envelope-rate |
quant | unbacked | — | PRE-REGISTERED 2026-08-22 (road-to-subagent-envelope-adoption Phase 1.3). BASELINE, measured before the pointer landed: **0 valid envelopes of 1,845 post-split stops**, window 2026-08-13T21:19:46Z through 2026-08-22T11:43:05Z, over agents/runtime/state/subagent-ledger/2026-08.jsonl. Verdict breakdown no_envelope1,817 ·fail28 ·ok0 ·no_message0; a further 4,543 rows carry the retiredabsentvocabulary and are excluded rather than folded in. Agent-type composition(null)1,725 ·general-purpose92 ·Explore28 — the null majority is the start-to-stop join rate of roughly 8 in 100 recorded elsewhere, and it is why no stop can be attributed to a dispatcher. METRIC:okdivided by post-split stops, reported bysrc/scripts/report_envelope_rate.ts, which prints the rate, the window bounds, the stop count and the ledger path on one line. THRESHOLD FOR THE FIRST WINDOW: greater than zero and rising. Deliberately NOT a percentage — a first window held to a high bar would fail for reasons the measurement cannot separate from the pointer's own effect, so the only thing the first window can establish is that the pointer is readable at all. POWER: the baseline denominator is 1,845; a first window below ~100 stops distinguishes nothing. FALSIFICATION: (1) a rate that stays 0 has at least three causes — the pointer is unreadable, the dominant path was misidentified, or workers on that path never emit a final assistant message — and a single rate cannot separate them, so a flat rate is reported as unresolved rather than as the pointer having failed; (2) the ledger is gitignored and machine-local, so every reading is one machine's drain traffic and no rate from it generalises; (3) a rate that rises without a contemporaneous pointer-removed arm does not establish that the pointer caused it — the historical baseline is temporally and compositionally confounded and both council seats refused it as a control. |
suggestion-capture-rate |
quant | unbacked | — | PRE-REGISTERED 2026-08-24 (road-to-suggestion-block-capture Phase 1.2), BEFORE any capture code exists. BASELINE: the model-carried comparator is resolved and dead — orchestration_recordcaptured 1 of 369 dispatches, and that figure "may not be cited for either direction" per its own entry, so it is the reason this instrument exists rather than a number this claim beats. THIS INSTRUMENT'S BASELINE IS ZERO BY CONSTRUCTION: the sinkagents/runtime/state/audit/suggestion-capture.jsonldoes not exist before the hook lands. METRIC: lines written to that sink divided by suggestion blocks emitted, where the denominator has a reading INDEPENDENT of the instrument under test — a contemporaneous emission log kept by the maintainer during the window. Without that independent denominator the instrument measures only itself and no rate is claimable. WINDOW: fourteen days, fixed insrc/config/suggestion-capture.json before any capture code ran, because a window whose length is chosen after the numbers are in is a window chosen to produce a number. THRESHOLD FOR THE FIRST WINDOW: greater than zero and rising. Deliberately NOT a rate figure — a first window held to a high bar fails for reasons the measurement cannot separate from the instrument's own readability, so the only thing it can establish is that capture happens at all. FALSIFICATION: a window in which the maintainer's log records blocks emitted and the sink carries zero lines DROPS this claim and parks the three consumer roadmaps' resume conditions as unsatisfiable by this instrument. SCOPE, measured rather than assumed: the payload probe covered Claude Code only (agents/evidence/analysis/suggestion-capture-probe.md), so any figure is one host on one machine and generalises to neither the other five bound platforms nor another operator's traffic. CLOCK CORRECTION 2026-08-24: no observation made before this date is admissible, and the fourteen days start at verified deployment of the FIXED instrument rather than at its commit. Until 2026-08-24 the concern declared main(now: Date = new Date())while the dispatcher callsmain(argv), so now.getTime()threw on every live turn and the instrument's own catch swallowed it — exit 0, no output, indistinguishable from a disabled hook. The consequence for THIS claim is specific rather than cosmetic: the falsification clause above drops the claim on a window whose sink carries zero lines, and applied to the broken period it would have dropped it for the wrong reason and parked three consumer roadmaps as permanently unsatisfiable on the strength of a bug.status: unbacked records that there is no evidence; it does not record which observations are admissible, which is why this correction is written here and not left to the reader. AI council 2/2. |
surgical-uninstall |
qual | backed | — | exec:vitest run tests/lib/json_pointers.test.ts tests/lib/host_hook_merge.test.ts -> 0 |
team-defect-finding-null |
quant | backed | — | internal/bench/reports/defect-finding.json#honest_null |
thin-inject-delivery-equivalence |
quant | backed | — | internal/bench/reports/thin-inject-2026-09-08.md#endpoints |
touched-file-quality-shadow-to-warn-bar |
quant | unbacked | — | PRE-REGISTERED 2026-10-07 (road-to-touched-file-quality-that-says-when-it-did-not-lookstep 4.1), BEFORE any reading is taken against it. The only reading that exists isagents/evidence/analysis/touched-file-quality-readings-2026-Q4.md, a replay of 30 merged commits, and that corpus is EXCLUDED below — so no number here can have been fitted to a result. |
turnaround-blocking-by-cause-targets |
quant | unbacked | — | PRE-REGISTERED 2026-10-07 (road-to-blocking-time-by-causestep 3.1), committed BEFORE either mitigation lands; the step order in the roadmap and the commit order here are the pre-registration. Reading:agents/evidence/analysis/turnaround-blocking-by-cause-2026-10.md (145 blocking calls, 802 blocking minutes over ten sessions). |
ui-conformance-behavioral-catch |
quant | backed | — | tests/scripts/ui_conformance_probe.test.ts#finds four of four planted defects |
unattended-demotion-gate |
quant | resolved-null | — | PRE-REGISTERED 2026-08-19 (road-to-long-horizon-execution Phase 4.2, sequencing UOTL Phase 7.3). REGISTERED BEFORE THE CAPABILITY EXISTS, which is the point and is stated rather than implied: at registration time NO unattended run has occurred and none can, because the headless spawn is deliberately unbuilt (unattended_guard.ts§ "Why the spawn is not in this file") and the budget defaults to both ceilings zero, which disables the lane rather than permitting it. So this entry cannot have been written around a number already in hand. THRESHOLD, fixed now: the rework rate of PRs produced by unattended runs, measured over the 14 days after each merges, must not exceed the same-window rate for attended PRs; a breach flipsmax_usd/max_tokensback to 0 in the same change that reports the number. REWORK is defined before any data exists, because a metric defined after the fact is chosen: a follow-up commit touching a file the run's PR touched, within 14 days of merge, excluding (a) commits by the same run continuing planned roadmap work, (b) pure dependency bumps, (c) reverts of an unrelated change that merely collide. POWER: at least 10 unattended PRs and 10 attended PRs in the comparison window, else UNDERPOWERED and no claim either way — a two-PR sample producing a favourable ratio is the failure this line exists against. FALSIFICATION: (1) an unattended PR that a human had to substantially rewrite counts as rework even when no commit touched the same file, and that case is recorded by hand rather than dropped because the mechanical definition missed it; (2) the comparison is rate-vs-rate, never absolute counts, since the two populations will not be the same size; (3) a rate that is LOWER for unattended runs is reported as-is and is NOT used to argue for widening the lane — this gate can close a lane, never open one. HONEST-NULL PATH: if the lane never runs (the spawn stays unbuilt, or the budget stays at zero), the recorded finding is that the gate was pre-registered and never had data, and the capability is closed rather than left indefinitely pending — the same D-5 shape this roadmap opens by naming. **HONEST-NULL PATH TAKEN 2026-08-19 — the SAME DAY it was registered, and that is stated rather than rounded** (R2 round 1, finding 2 caught this entry claiming "one day after"; registration above is dated 2026-08-19 too, so the elapsed time is zero). A pre-registration closed on its own registration date deserves the suspicion it attracts, so here is why it is not a threshold written around a result: the entry was registered when the spawn was DEFERRED, and closed when the spawn became REFUSED. Nothing measured moved in between — a decision did, and it is recorded with its council and its reasoning. The threshold never met data in either state. The lane will not run: the headless spawn is no longer "deliberately unbuilt" but a published refusal (road-to-long-horizon-execution 4.0, AI council 2026-08-19), and the two acceptance criteria that depended on it are cancelled WILL-NOT-MEASURE. So the recorded finding is exactly the one this path fixed in advance: **the gate was pre-registered, never had data, and the capability is closed** — zero unattended PRs against the ≥ 10-vs-10 power floor, which is not a null result but an absent population, and the entry says which. Nothing here is read as evidence in either direction, and specifically not as "unattended runs are safe": an unrun lane has no rework rate. The threshold, the rework definition and the power floor are left EXACTLY as registered so that a future roadmap reopening the capability inherits a bar written before anyone knew the answer; reopening requires that new roadmap, never a re-reading of this entry. The reopen trigger is 4.0's: the first checkpoint written by a real dying run. STATUSresolved-null, not unbacked (R2 round 2, finding 3): the ledger already has the terminal status for exactly this state, and leaving a closed question inside the documented-debt inventory would inflate the unbacked count and leave it looking pending — the indefinite-pending shape this entry argues against, reproduced by the entry that argues against it. |
user-out-of-loop-baseline |
quant | unbacked | — | PRE-REGISTERED 2026-08-17 (road-to-user-out-of-the-loop Phase 0 Step 3 — registered BEFORE the instrument had recorded a single observation, and deliberately BEFORE any Phase 1 mechanism ships, so the baseline cannot be read after the change it is meant to judge). INSTRUMENT: the interruption-ledger concern (src/scripts/hooks/interruption_ledger_hook.ts, stopslot, capture-only) writes one{run_id, turn, kind, class, roadmap}line per turn toagents/runtime/state/interruptions.jsonl; src/scripts/interruption_report.ts reads it. DEFINITION fixed before data, and it is three classes rather than two on purpose: a CONTACT is a turn whose closing paragraph either ends in a question (ask) or yields the decision without one (handback). Counting only ?would score this package's own preferred hand-back shape as zero contacts and make the metric flatter the design — the Risk-6 failure the roadmap names. POWER CAVEAT recorded at registration, not discovered later: the rolling chat history is a buffer, not an archive — measured the day this was written it held **5 sessions, all from one day**, against the 30-session conformance window the step asks for. The report therefore reportssessions_foundnext towindow_requestedand flagswindow_short; a number computed over a short window may not be cited as an N-session baseline in either direction. FALSIFICATION fixed before data: (1) the QUALITY ANCHOR is the held defect rate — a contact reduction that moves the defect rate is a FAIL regardless of its own number, and the two are never reported apart; (2) the baseline must rest on ≥ 20 recorded runs before any post-change comparison is made, and a comparison against fewer is reported UNDERPOWERED, never as a win; (3) a run present in the ledger but not the history, or the reverse, is reported with the missing axis null and is never scored as zero — scoring an unmeasured run as zero contacts is the arithmetic that would manufacture the result. HONEST-NULL PATH: if the median does not move, or moves only with the defect rate, the null is published and Phase 1's mechanisms are judged on it rather than the claim being re-scoped to a family that happened to win. |
utilization-window-decidability |
comparative | unbacked | — | PRE-REGISTERED 2026-07-12 (road-to-feedback-8.11-2 Phase 0 — no goalpost-moving after the numbers land; criteria at docs/design/utilization-window-criteria.md). Floor fixed BEFORE data: >=100 task boundaries AND >=2 hosts (or the documented degraded form) AND >=45 elapsed days; decision rules D1 (loaded-never-consulted -> retirement-candidate list), D2 (consulted-never-applied <10% applied-ratio at >=5 consultations -> trigger-review queue), D3 (above floor -> >=1 named decision per kind or a recorded why-not), D4 (below floor after one extension -> honest null, lifecycle/ledger gates stay closed). Kernel + safety floors exempt by construction. |
web-launch-readiness-finds-more |
quant | unbacked | — | PRE-REGISTERED 2026-08-25 (road-to-web-launch-readiness Phase 0.4), BEFORE any skill code exists -- Phase 0's exit condition is "zero skill code written", which is what makes this a pre-registration rather than a description of what the skill turned out to do. BASELINE: zero, by construction. The estate has no web-surface coverage to beat: grep -riloversrc/skills/ src/rules/ src/domains/returns 0 files forrobots, 0 for noindex, 0 for meta descriptionand 0 forlighthouse, and the nearest-named skill (src/skills/launch-readiness/SKILL.md, 220 lines) scores 0 on all eight of 404 / robots / noindex / canonical / meta / alt text / analytics / legal. Re-measured at 4014008f7 on 2026-08-25; full note at agents/evidence/analysis/web-surface-coverage-g0-2026-08-25.md. DESIGN: three fixture sites of KNOWN defect state -- one local-business-shaped static site, one SaaS-shaped app, one docs site -- seeded before either arm runs with a staging noindex, a missing custom 404, missing metadata on two routes, three images without alternative text, and a missing privacy-page link. COMPARATOR: identical model, bare audit prompt, same fixtures, same run. METRICS: precision and recall against the ground-truth list. HARD GATE, and it is the reason this claim can fail while scoring well: one site-type-IRRELEVANT decoy is seeded -- a missing team photo on the SaaS app -- and flagging it is a classification failure that DROPS this claim REGARDLESS of recall. A skill that finds everything by flagging everything is the failure mode a recall threshold cannot see, which is why the decoy is a gate and not a metric. FALSIFICATION: (1) the decoy is flagged; (2) recall does not exceed the comparator arm's; (3) the fixtures cannot be built to a ground truth both arms are scored against, in which case the claim is UNDERPOWERED rather than dropped -- an unbuildable fixture says nothing about the skill. ON DROP: the skill does not ship enabled, and road-to-web-launch-readiness records the null as its outcome. SCOPE: three fixtures on one model is one measurement, not a general result about audit skills; it establishes whether THIS skill beats a bare prompt on THESE defects and generalises to neither another model nor another defect set. |
wedge-hollow-detection |
quant | backed | — | internal/bench/orchestration/pv-a3-results.md#token_delta_provenance: measured |
worker-capsule-trigger-arm |
comparative | unbacked | — | PRE-REGISTERED 2026-08-09 (road-to-worker-generation-recycling Phase 1.4 — registered BEFORE the first shadow capsule is read; the mechanism ships shadow-only, so no capsule has been scored at registration time). CAPSULE-QUALITY RUBRIC, fixed here, five binary criteria scored 0-5 per capsule — (1) remaining[]names every open item the task still needs, no silent drops; (2)decisions[]names each choice a successor would otherwise silently re-open; (3)assumptions[]is non-empty and every entry carries a resolvingbasisref; (4) everydone[] ref resolves to a real file/line; (5) a successor briefed on the ORIGINAL brief plus the capsule alone takes a first action that neither repeats completed work nor asks for a re-brief. ADOPTION MARGIN, fixed BEFORE data: an arm is adopted only if, on paired samples from the same runs, it fires at a median of >= 2 steps earlier AND its capsules score >= 4/5 on the rubric with no regression against the other arm; an arm that wins on earliness while dropping below 4/5 is NOT adopted, because an earlier bad capsule is worse than a later good one. Sample floor: >= 30 shadow capsules with BOTH trigger points recorded (watermark_step, saturation_step, trigger_arm_earlieron theorchestration_recordline). Instrument:src/scripts/_lib/capsule_trigger.ts (compareTriggers, earlierArm), term-frequency only, no embeddings. HONEST-NULL consequence, pre-authorised: BOTH arms losing (neither reaches 4/5, or the margin is not met) is a publishable result that closes the mechanism as default-off — it is the expected-value outcome given the standing orchestration-observed-dispatch-cost null, and it must be cheap to record. Token delta is reported as a pair with quality and is explicitly NOT the claim. |
5. Verify it yourself
Section titled “5. Verify it yourself”On a fresh checkout, reproduce the claims above:
task check-claims # every markered public claim binds to resolvable evidencetask check-refs # no broken internal referencestask check-skill-gaps # every logged known-limit cites a real witness testtask check-comparison # every comparison-table "our evidence" pointer resolves./scripts-run src/scripts/skill_eval_coverage # behavioural-eval coverage, per tier./scripts-run src/scripts/skill_eval_coverage --check # the ratchet: coverage may not droptask build-proof-check # this page is in sync with its sourcesIf a claim ever loses its binding, or this page drifts from the ledger, CI goes red. Reproducibility is the proof.