Verify it yourself
Proof — verify our claims yourself
Section titled “Proof — verify our claims yourself”We sell falsifiability, so the selling is machine-checked. Every public claim binds to resolvable evidence; a skeptic can reproduce every check below on a fresh checkout. This page is itself generated from those sources and fails CI if it drifts.
What this prevents
Section titled “What this prevents”Each row is a failure mode, not a feature — and each cites the ledger claim that backs it. A row whose claim stopped resolving would fail this page’s own generator, so the table cannot outlive its evidence.
| Failure this prevents | How | Backed by |
|---|---|---|
| A rule says MUST, and nothing anywhere enforces it — so the rule reads as a guarantee while being honour-system. | Each rule declares enforced_by:, and the check resolves it: a validator reachable from no taskfile, workflow, or hook manifest counts as unwired, not as covered. Undeclared counts as uncovered. |
enforcement-coverage-resolved ⟳ |
| A shipped artefact carries a hidden-Unicode or instruction-smuggling payload, and reaches consumers because only the source tree was scanned. | Source and the condensed projection are scanned in CI; a finding blocks the release before publish, not merely the merge. | shipped-artifacts-hidden-instruction-scanned ⟳ |
| A count in public prose drifts from the source, and the marketing number quietly stops being true. | Counts are generated from source and drift-checked; the build fails on a count-shaped prose mention that disagrees, or on two different numbers for the same artefact kind. | skill-count ⟳ |
| A host-coverage claim understates or overstates what actually works — for months, on the one surface everybody reads. | The number is pinned by a test over real detection, and the test is re-run as the claim evidence. | host-agent-count ⟳ |
| A non-coding domain skill implies proven correctness it never had, because it was forged on a different domain and never checked against one. | Domain skills are labelled unvalidated until they pass a sourced domain-truth fixture, and the validated count is ratcheted so it cannot quietly fall. | domain-soundness-scoped ⟳ |
| Behavioural-eval coverage regresses as skills are added, and the suite looks healthy because only the absolute number is reported. | Coverage is measured per tier and ratcheted in CI, so it can only rise; the gap is published rather than implied away. | eval-coverage-ratcheted ⟳ |
⟳ — the claim carries exec: evidence: CI re-runs the command and
compares its exit code, so a stale row turns the build red rather than
ageing quietly into marketing.
See it run (< 60s, real output)
Section titled “See it run (< 60s, real output)”
The recording is of the exact commands in § 5, nothing staged. Its
commands are re-executed in CI (.github/workflows/proof-demo.yml), so a
demo that showed something broken would turn CI red — the recording
cannot drift from current behavior.
1. Every public claim binds to evidence
Section titled “1. Every public claim binds to evidence”Rendered from docs/CLAIMS.md. A <!-- claim:ID --> marker in
README/docs must resolve to a backed entry here with a resolving
evidence pointer, or task check-claims fails the build.
| Claim | Kind | Evidence | Resolves |
|---|---|---|---|
| The two surfaces the provider-lifecycle contract obliges to agree on an adapter tier — the adapter header and the xml example — do agree, and the surface that read stale is no longer hand-written: § 5 is generated, so the drift this claim was opened over cannot recur. | qual | docs/contracts/provider-lifecycle.md#Current tier assignment |
✅ |
adversarial-review structures its critique as branching exploration with explicit pruning rather than a single pass, following the Tree-of-Thoughts formulation. |
qual | https://arxiv.org/abs/2305.10601 (2026-08-11) |
✅ |
The .augment-plugin/ manifest version is the package version, not an independent plugin-API version, and every version-bearing file the release workflow triggers on is read by a job in that workflow. |
qual | src/scripts/lint_marketplace.ts#check_augment_manifests |
✅ |
analysis-autonomous-mode runs iterative self-critique between steps rather than only at the end, following the Self-Refine formulation. |
qual | https://arxiv.org/abs/2303.17651 (2026-08-11) |
✅ |
bug-analyzer verifies each candidate root cause against a concrete trigger before reporting it, following the Chain-of-Verification formulation — which is the mechanism behind its “never invent issues” constraint. |
qual | https://arxiv.org/abs/2309.11495 (2026-08-11) |
✅ |
| The release process is documented as an inheritable runbook + succession doc, and the project’s bus-factor (trailing-90-day distinct human reviewers) is tracked and reported truthfully — currently 1, not implied to be more. The doc separately reports the distinct MERGER count (2, one of them an unreviewed self-merge) so the reviewer figure cannot be inflated by conflating the two. | qual | docs/succession.md#trailing 90 days |
✅ |
| On the pre-registered 2-arm retrieval benchmark (18 hand-verified code-structure questions across 3 real repos, ground truth hash-bound before the run, deterministic, zero model calls), the native code graph scored mean recall 0.365 vs grep 0.797 on the 15 graph-shaped questions (delta -43.2 pp against a pre-declared +10 pp win threshold) and 0.111 vs 0.833 on the negative controls. HONEST NULL — measured root cause: TS arrow-function exports produce no symbol nodes (170 TS vs 13,428 PHP symbol nodes on same-shaped repos) and string-keyed dynamic consumers have no static edge. Consequence bound: code_graph.enabled stays false permanently; deprecation at the next major, removal the major after unless external evidence appears. | quant | internal/bench/reports/code-graph-vs-grep.md#Verdict — NULL, decisively |
✅ |
| 202 commands. | quant | exec:check_artefact_count_messaging -> 0 |
✅ |
On the first post-fix behaviour-conformance measurement (round 5, /analyze:conformance --limit 30, run 2026-08-07, every violation split by its own timestamp against the 2026-08-06 carrier merge), both BLOCKING pre_tool_use guards eliminated their classes — unauthorized irreversible git ops 8 → 0, evaluator prompt pre-loading its verdict 1 → 0 — while neither advisory carrier did: language-mirror violations fell 555 → 19 under advisory state injection at user_prompt_submit (−96.6%, not zero) and verification-claimed-on-empty-output fell 4 → 1 at post_tool_use. Advisory reduces massively; only blocking eliminates. Scope bounds, inseparable from the numbers: the post-fix corpus is ONE session (~600 assistant turns), so this is a recorded prior, not a law; the language pin was verified PRESENT on the violating post-fix turns (which is why an UNBOUNDED re-pin — restating it on every tool call — stays refused; what shipped 2026-08-20 instead is a BOUNDED re-emit, once per 150 tool calls, against a red baseline this round-5 corpus could not produce and whose own revisit clause called for: 11 English replies to a German user, all of them 179+ tool calls past the pin, none below. Corrected here because the unqualified earlier wording — “higher injection frequency is refused as a fix” — is a claim the shipped code contradicts; the reasoning, the measured distances, and the accepted false-fire cost are in agents/settings/contexts/reminder-injection-verdict.md); and the 555 pre-merge count is partly contaminated by the synthetic-turn mis-pin fixed in the same round (isSyntheticPrompt). NOT PRE-REGISTERED, and it could not have been: this is a post-hoc audit reading taken after the carrier landed, so no bar existed to freeze before the data. What stands in for pre-registration is that the one choice a post-hoc split could game — where before ends and after begins — was not chosen by the measurement: the boundary is the carrier merge’s own commit timestamp (2026-08-06) and every violation is assigned by its own timestamp against it. The detector is deterministic and re-runnable over the same transcript store, so the numbers are reproducible rather than attested. Read it as a recorded prior, never as a pre-registered result; the sibling claim that DOES carry a frozen bar is the scoped-rule absence experiment, pre-registered precisely because its data does not exist yet. |
quant | src/domains/analysis-workbench/analyze/conformance/command.md#Both blocking carriers reached zero |
✅ |
A MEASURED-BUT-NOT-SHIPPED experiment — the thin rule projection reduced eager rule load 78,513 → 13,881 GPT-tokens (whole always-loaded projection 98,529 → 33,897, ~65.6%), but FAILED the quality gate (thin win-rate 36.2% vs required 48%) and does not ship; it un-defers only behind discipline_profile: essential. Shipped behavior does NOT include this reduction. Method: agent-config benchmark over the pinned token baseline; the baseline is the honest “what the user pays if everything loads eagerly”, NOT a synthetic full-corpus strawman (council Q4); quality gate per the Phase-0 paired judge run. |
quant | internal/bench/reports/token-baseline.json#eager_rule_load |
✅ |
An eligible mid-flight cli failure with a constructible api twin loses zero council seats WHEN THE PROJECTED-SPEND GATE PERMITS THE RETRY — the seat answers over the api rung instead of dropping out of the pass. A retry the budget REFUSES is outside the claim: the original failure stands, the seat is absent, and fallback_skipped: cost_budget says so. |
qual | exec:vitest run tests/scripts/ai_council/council_cli.test.ts -> 0 |
✅ |
A council member’s unparseable answer is separable from a member that found nothing — 2/7 parse_failed, 1/7 empty, 4/7 parsed over the seven recorded answers in tests/fixtures/council-parse-corpus/, which is a fixture-corpus denominator and NOT live traffic; reproduce with ./scripts-run src/scripts/council_parse_rate. |
quant | exec:vitest run tests/scripts/ai_council/parse_corpus.test.ts -> 0 |
✅ |
The first cross-vendor parity pass (5 orchestration-corpus tasks × 2 vendors × 3 repeats, identical prompts via the council transport, $0.16) measured real per-host finding-count differences — claude-sonnet-4-5 surfaced ~2× the findings of gpt-4o on the multi-file analysis task (median 11 vs 5) while both vendors were identical on the planted hollow-implementation task (2 vs 2) and perfectly silent on the clean-code negative control (0 vs 0, no spurious findings). The per-task finding_floor values are calibrated from the cross-host lower envelope and the gate is armed. |
quant | internal/bench/reports/parity-count.json#min over hosts of median |
✅ |
The scoped-projection default for new installs ships 228 of 299 skills (untagged core plus engineering/maintainer packs), an approximately 25% reduction of the skill-catalog surface (a reduction of 71 projected entries; the token figures measured 2026-07-27 at the then-283-skill catalog were about 577k to about 428k approximated tokens and are NOT rescaled here). Both figures are generated, not typed: reproduce them with ./scripts-run src/scripts/count_scoped_projection, which partitions the canonical skill catalog with the same predicate install.ts applies when it prunes a real tree. |
quant | exec:update_counts --check -> 0 |
✅ |
On a 32-file labelled clean-UI corpus, 18 of the 19 shipped design-slop rules produced zero false positives; the nineteenth (slop-c6-lock-colour, catalog C6) fired on 4 of the 32 files and was demoted to judgment-only rather than tuned. |
quant | internal/bench/corpora/design-slop-fp-PREREG.md#The ceiling, declared before the run |
✅ |
| On a weak host (claude-haiku-4-5) the package produces a significant, placebo-controlled discipline lift on scope/downstream traps; on a strong host the same measurement is a published null — the package transplants discipline a weak model lacks, not model intelligence. | quant | docs/benchmark.md#weak-host-specific |
✅ |
| The non-coding domain skills (finance/founder/ops/content) are forged on TS/PHP and labeled unvalidated until they pass a sourced domain-truth fixture; no public prose implies proven domain correctness, and the validated count is CI-ratcheted. | qual | exec:domain_soundness_status --check -> 0 |
✅ |
The validated non-coding domain-skill count is pinned and CI-ratcheted at a maintainer-set floor (9 of 20 default-surface skills carry a sourced evals/domain-truth.json fixture at pin time, 2026-07-11 — 5 deterministic, keys from cited formulas; 4 rubric, criteria matching a named external practice); the floor only rises via a maintainer --write-floor after a new sourced fixture lands. |
quant | exec:domain_soundness_status -> 0 |
✅ |
| On the READ-ONLY FAN-OUT slice family, tier-downshifted subagent dispatch (lite/haiku vs session-tier-proxy sonnet) nets a ≥30% USD-weighted token-cost reduction at held quality — measured 2026-07-08 (n=10 paired live dispatches, 20 telemetry lines): 10/10 exact-match on BOTH arms, 29.4% fewer raw tokens, 76.5% USD-weighted cost reduction at the 3x haiku↔sonnet price ratio. FAMILY-SCOPED — the mechanical-edit family is unmeasured and its downshift (incl. the deferred tier downgrades of existing units) stays gated. Negative control held: an open-ended synthesis/unknown slice never resolves below the session tier (inferSliceTier → medium/inherit, never lite). | quant | internal/bench/routing-downshift/results-2026-07-08.md#FAMILY-SCOPED PROVE |
✅ |
The text-layer-only boundary above is machine-enforced, not asserted: a scope-guard test fails if any corpus entry declares a non-text layer, and that guard is itself falsified by a test that splices in a layer: file PNG-metadata fixture and requires the guard to fail. The corpus is sha256-frozen and was committed BEFORE any detector existed, so no detector was tuned against it. |
qual | tests/scripts/encoding_corpus.test.ts#FAILS when a deliberately out-of-scope fixture is added |
✅ |
| The retrieval sanitize floor covers the TEXT layer only, by construction — never file or network channels (image / audio / PDF / DNS / TCP / file-metadata steganography) and never semantic evasion (word choice, phrasing, garden-path constructions, word-order permutation). On the frozen 653-entry corpus it strips or flags 99.00% of in-scope positives (100.00% on the unambiguous zero-width / bidi / variation-selector classes) at a 0.00% false-positive rate over 353 real in-repo negatives, with zero added model spend and 0.018 ms p95 per message. Exactly ONE of the seven added channels removes bytes; the other six report and pass the text through unchanged — this is NOT a claim to block steganography. | quant | agents/evidence/reports/encoding-floor-measurement.md#Selected branch: ADOPT |
✅ |
The share of governed rules carrying a backstop that fails a CI build is RESOLVED, not declared, and is published in exactly ONE place — docs/proof.md § 4b, projected from check_enforcement_coverage, which prints its denominator together with the frame that produced it (internal/reports/enforcement-coverage.json is that same output on disk). No figure is restated in this entry, deliberately, and check_enforcement_denominator reds when one appears in a published doc the resolver did not generate: until 2026-08-23 this tree carried FIVE different numbers for the one property, and every previous correction fixed a figure while leaving the plurality intact — a number that is right today is how it came back each time. Resolution means a declared validator: counts only when the script exists AND a GITHUB WORKFLOW reaches it (transitively, so a sub-check under a wired umbrella counts), while a hook registered fail_closed: false resolves to observer, never validator. That distinction is not cosmetic: it once let “named in a taskfile” read as “fails the build” while NO workflow invoked task ci, ci-strict, or ci-fast, so most of the counted validators only ran when a human typed the command. Splitting validator (CI runs it) from validator-local (only a human does) cut the honest figure to roughly a third at the time; wiring the taskfile-only gates into rule-backstops.yml restored it, this time meaning what the headline says, and local_only is ratcheted at zero so a gate cannot drop back out silently. Wiring them also surfaced that five were failing invisibly, 37 findings deep; those are cleared and the baseline in rule-backstop-debt.json stands at zero, so the ratchet enforces rather than tolerates. Roughly two thirds of the 37 were never violations — the gates were misreading allowances their own rules already grant (license-required attribution, multi-stack peer examples), which is the same class of defect one level down. The ratio has not risen across five releases, and the reason is worth stating rather than reading as stagnation: the ratchet prevents regression, it does not raise the level. An undeclared rule counts as uncovered, never excluded — an honest recorded gap beats a false claim of coverage — including the two scale/history pack rules that ship enforced by lint_persistence in consumer CI, which this resolver, scoped to THIS repo’s workflows, correctly does not count. |
quant | exec:check_enforcement_coverage --check -> 0 |
✅ |
| Exactly one enforcement denominator is quotable, it names the frame that produced it, and no published doc restates it by hand. | quant | exec:check_enforcement_denominator -> 0 |
✅ |
| The lift-carrying essential cut (kernel + downstream-changes) keeps a significant weak-host discipline lift at a fraction of the full load’s tokens, and the lift is FAMILY- and HOST-SCOPED — measured on three hosts: claude-haiku-4-5 (weak) shows the family-scoped lift (trapE 0.533→1.000, 7/7 discordant, corpus cost 1.71x); claude-sonnet-4-6 (strong) is a ceiling null; gpt-5-mini (non-Claude weak, codex prompt-prepend surface) FAILED replication with headroom (corpus Δ=+0.024 p=0.70, capability trend n.s. — no harm claimed, injection-surface confound documented). Therefore discipline_profile: auto enables the lift only where measured (vendor-granular unknown_defaults). Non-claims — the balanced router profile was removed after a NULL measurement (p=0.81, n=24); no full-tier recommendation exists; no cross-vendor lift is claimed. | quant | docs/benchmark.md#REPLICATION FAILED |
✅ |
| Behavioural-eval coverage is measured per tier and CI-ratcheted so it can only rise; the current coverage and its gap are published, never implied as “264 evaluated skills”. | qual | exec:skill_eval_coverage --check -> 0 |
✅ |
| A credential-free prescription layer reads content the host’s own web tools cannot fetch at all — Reddit thread text (Atom feeds) AND comment ranking plus reply nesting (server-rendered HTML), and a single named tweet (the platform’s own oEmbed endpoint). Measured 2026-07-25 from a residential network against a pre-registered 6-task-per-channel set with a native control and thresholds frozen before the run: reddit tier 1 6/6, reddit tier 2 6/6, twitter-oembed 6/6, native 0/6 on both Reddit tiers. Zero credentials, zero resident processes, zero auto-installs. Scope bounds that travel WITH the claim: (a) the twitter gap is narrower than 6/6 suggests — native also passed 2 of those 6 via third-party mirrors, so the channel earns its place only on tweets nothing mirrors; (b) reddit tier 2 is on an announced closing path and ships with a kill-switch keyed on an OBSERVED login wall; (c) youtube-transcripts is PARKED, not shipped — its backend is human-installed by contract and was never exercised; (d) residential network is load-bearing, and CI is explicitly not a bench environment. | quant | docs/benchmark.md#ship-gated-reach |
✅ |
Four adjacent governance properties were closed as regression tests rather than phases, on the expectation that each was already true. One of the four was. (a) enforcement never branches on a base-model refusal string — holds; the single module compiling refusal regexes only ever escalates, and no refusal branch reaches an allow decision. (b) a capability gate resolves only from trusted config — VIOLATED: the runtime dispatcher returned ready for a skill whose own frontmatter declared safety_mode: strict and granted itself 2 tools absent from the 2-entry registry, while the validator implementing that allowlist had zero production callers; now wired. (c) caller-agnosticism — holds: 0 caller-identity inputs reach a gate verdict, and the 3 platforms that carry the blocking slot are pinned so none silently loses a concern. (d) constraint monotonicity — holds: 0 blocking gates read persisted state, with 1 advisory anti-nag exception named in the test. All 7 inverted properties produced a failing test. |
quant | internal/bench/reports/governance-invariants.json#adjacent_properties |
✅ |
The council aggregation cannot be steered against a refusal. Pre-registered spike S0.1 measured that it WAS classification-steerable — w_total counted only members whose stance line parsed, so a refusal phrased as prose left the quorum and made consensus easier: steering margin 0.6667 (margin -0.25 parsed vs +0.4167 unparsed) with the outcome flipping from no-consensus to Adopt. High severity because the direction was the dangerous one. Fixed in the same change: a member who responded counts toward the quorum whether or not its stance parsed, and the post-fix steering margin is exactly 0. The divergence signal is an observation and is asserted never to reach the scoring path. No observed attack prompted this; the expected outcome was a null and it was pre-registered as such before the run. |
quant | internal/bench/reports/governance-invariants.json#s0_1_aggregation_steerability |
✅ |
Pre-registered spike S0.2 measured that both of this package’s fail-closed gates judged the shape of ONE action rather than the effect, so a sequence whose every step they allowed composed into the outcome they exist to prevent — leak count 2 of 2 gated outcomes, with all 4 single-step controls blocking correctly. Both were moved to the effect they govern. The positive control is check_secret_leak, which scopes to the cumulative diff against a base ref and therefore returned null before any fix — decomposition gains nothing against an effect-scoped gate. Scope bound, published rather than swept: mv / chmod / rm against .git/hooks/* still reach the first outcome and are asserted as an open gap, because recognising them would make a fail-closed guard a shell sandbox. Decomposition in this package is model-carried — there is no executable subagent dispatcher — so the measured layer is the PreToolUse hook layer; the 2026-08-02 council dissent on whether that is a faithful discharge is recorded in the spike header. |
quant | internal/bench/reports/governance-invariants.json#s0_2_decomposition_laundering |
✅ |
| Pre-registered spike S0.3 measured that a stated uncertainty, hedge, or provenance marker survives this package’s telegraph condenser into the audited text — 3 marker classes, 10 fixture cases, marker-loss count 0, negation count preserved. Honest null, published with the spike wired as the regression test. The first run’s two “failures” were fixture defects (carriers written as phrases containing an article the condenser is documented to drop) and are recorded as an unmet premise rather than edited away. Scope bound: this protects a marker the agent DID emit; it cannot make an agent state an uncertainty it never stated. | quant | internal/bench/reports/governance-invariants.json#s0_3_marker_survival |
✅ |
Hook dispatch runs as one precompiled node process with all concerns in-process, and the CI latency gate measures the REAL invocation path — the exact command hooks.json installs (bash wrapper + install-shape probes + dispatcher), via bench_hook_latency --gate --via-cli, whose per-event commands come from the same generator that writes hooks.json — not the bare bundle the pre-repair gate measured. The repair (road-to-hook-latency-repair) is pinned before/after on one machine in the committed baseline history: pre-fix CLI path pre_tool_use p95 164 ms → post-fix bundle-direct path p95 84 ms (darwin dev, warm cache, n=50 each; the pre-fix path measured ~450–500 ms/event on a 1-vCPU container). Budget re-derived 2026-08-19: pre_tool_use p95 175 ms (was 150), any event 250 ms on GitHub-hosted CI runners, both BLOCKING. The raise is measured rather than bent around a red check — 150 sat INSIDE the observed legitimate distribution (green main runs 115, 120, 146, 150, 150 ms; reds 151, 152, 152, 154 ms; a comment-only PR run at 157 ms) and failed 5 of 33 tests.yml runs on main over 2026-08-16..18, i.e. 15%, with hand re-runs as the standing workaround. The budget file records that distribution, the failure rate, the cited derivation rule and a revisit-if that routes the NEXT breach to the cause rather than to a third raise: a GREEN run whose p50 rises above 160 ms is the cost growing, not the runner. Down from ~1.6 s p50 on the retired CLI-to-bash-to-tsx per-concern-respawn chain. |
quant | docs/hook-latency.json#invocation_path |
✅ |
23 host agents are detected and inventoried; 20 receive a written config surface (18 projection + 1 plugin + 1 bundle target) and 3 are export-only (aider, zed, jetbrains). The count is enforced, not asserted — knownToolIds() is pinned at 23 by a test whose assertion literal IS the number, and src/config/surface-matrix.yml is held in set-equality with the installer’s own user-scope path map by lint_surface_matrix, so a host added to one and not the other fails the build. This entry stood unbacked while naming its own unblocking condition (“once surface-matrix.yml exists, bind the count to that file and flip”); the condition was met and nothing fired, so the shipped figure stayed “7+” — understating real coverage by 3x. |
quant | exec:vitest run tests/install/toolDetection.test.ts -> 0 |
✅ |
| On the fixture corpus (n = 20 before/after pairs, 16 length-controlled within ±25%), the humanizer pass removes every mechanically detected AI-writing tell (mean hard hits 0.9 → 0, cluster score 53.97 → 0 per 500 words, dash density 9.22 → 0), and a blind judge (claude-sonnet-4-5, deterministic per-pair A/B seed) preferred the humanized text in 16/16 length-controlled pairs. Scope note — the “before” fixtures were deliberately tell-seeded, so this measures seeded-tell removal on a self-constructed corpus, NOT real-draft improvement; real-world lift is unmeasured until step 4b has processed real ghostwriter drafts (see the road-to-humanizer-hardening live-usage blocker). | quant | internal/bench/reports/humanizer-v1.md#prefers the humanized text |
✅ |
The injection-scan post_tool_use detector measures 99.00% recall (297/300) and a 0.85% false-positive rate (3/353) over the frozen corpus at internal/bench/corpora/encoding-channels/. Both figures are properties of THAT corpus and not of tool output in the wild — the evidence report says so itself: web fetches and MCP responses are a different distribution, so the 0.85% predicts nothing about a repository whose tool output is largely prose about prompt injection. Distinct from encoding-floor-text-layer-only, which measures the stripping pipeline at a 0.00% false-positive rate over a different scope; the number here belongs to the reporting half, which had never been measured separately before. The detector ships enabled: false and warn-only, so neither figure is a claim that anything is blocked. |
quant | agents/evidence/reports/injection-detector-wiring.md#The numbers |
✅ |
| A fresh registry install of this package carries zero high/critical npm-audit findings on the runtime dependency tree (0 vulnerabilities total at last verification), gated on every PR and every release PR. | quant | .github/workflows/release-validation.yml#npm audit --omit=dev --audit-level=high |
✅ |
The judge-skill family (judge-bug-hunter, judge-code-quality, judge-security-auditor, judge-synthesis, judge-test-coverage, and the /review-changes dispatcher) implements the LLM-as-a-judge pattern — a specialized model scoring another model’s output against a rubric — and names position bias and self-consistency as its known failure modes. SCOPE: the pointer backs the pattern and its named failure modes, which is what the cited work establishes. It does NOT back any claim about how this repo mitigates them; a statement about the dispatcher’s own behaviour needs a repo pointer, not a paper. |
qual | https://arxiv.org/abs/2306.05685 (2026-08-11) |
✅ |
| Word-boundary-anchored keyword matching reduced unintended rule activations by 12.5% (495 → 433) over the 302-prompt matrix-derived corpus with zero intended positives lost. Disclosure — the derived corpus was co-edited in the same change (6 German positives re-authored to standalone tokens; verb-inflection recall is a documented accepted cost). The circularity is broken by an independent replay over 49 UN-edited real-corpus prompts: recall 15/17 in BOTH arms (zero labels lost to anchoring), unintended activations 110 → 99 (−10.0%). | quant | agents/evidence/analysis/anchoring-independent-replay-2026-08.md |
✅ |
NONE of the backed ledger claims are machine-re-verifiable today — every evidence pointer is checked for existence (file present, substring present, URL carries a date), never for truth, so a claim pointing at a stale artefact stays backed indefinitely. A measured minority COULD carry a re-executing exec: form, clearing the >= 10 pp threshold that was pre-registered before the count was taken, which is why that form is scheduled rather than assumed. The rest cannot: paid or stochastic benchmark runs no CI job can re-derive, and prose contracts. Exact counts live in the evidence file and are NOT restated here on purpose — this entry hard-coded its denominator twice and drifted within a day both times (25 when the ledger held 26, then 26 when it held 27) while CI stayed green, because the pointer resolved. check_claims now compares the stored denominator against the live ledger and fails on divergence. A number a human retypes on every ledger edit will drift; the fix was to stop retyping it. |
quant | internal/reports/exec-evidence-feasibility.json#"exec_feasible" |
✅ |
A hand-rolled, dependency-free BM25 + trigram lexical index resolves the “recalls but does not rank” gap: on the retrieval-precision corpus (9 keyword-overlapping-confuser tasks) it drives the mean top tie-set from 3.333 (the _score bucket scorer) to 1.0 — every needed decision uniquely top-ranked — with precision@1 and precision@5 unchanged at 1.0. Method: deterministic, model-free re-ranking of the SAME retrieved entry set; both scorers measured over the identical store via measure_lexical_ranking.ts. Cross-artefact note (2026-07-25): this baseline (3.333) is the value in THIS claim’s own artefact and is cited correctly, but the retrieval-precision artefact records 4.111 for the same scorer on the same corpus — two bench scripts disagree. The lift direction (ties collapse to 1.0) holds under either baseline; the discrepancy itself is unresolved and recorded rather than smoothed. |
quant | exec:measure_lexical_ranking -> 0 |
✅ |
| Registering an MCP server with this package costs standing context on every session, and the kernel server’s 20-tool surface costs 3,886 tokens of it while the two-tool lite surface is capped at 600. | quant | agents/evidence/metrics/mcp-tool-standing-cost.jsonl#tool_search_threshold |
✅ |
| The whole layer is compiled into host agents with zero runtime daemon. | qual | docs/contracts/no-runtime-boundary.md#file-first, no-runtime suite |
✅ |
| On a 12-fixture option-decision corpus (3 arms × 2 providers, blind rubric judge claude-opus-4-8, pre-registered hypotheses), famous-figure identity framing added nothing beyond the underlying method text (method 5.04 vs figure 4.88, Δ=0.17, sign-test p=0.607), and provider diversity moved judged quality ~15× more than persona identity (provider Δ=2.58 vs identity Δ=0.17); the whole persona layer lifted only +0.08 over bare prompts. Honest null — persona panel-mode stays CUT, evidence-closed. | quant | internal/bench/reports/persona-placebo.json#honest-null |
✅ |
| We publish our own measured null results and retire or constrain features when the evidence does not support them. Deliberately falsifiable — every published null links the run that produced it; find one that does not resolve and this line updates. | qual | docs/benchmark.md#honest |
✅ |
On the live end-to-end retrieval benchmark (9 tasks × 3 arms × 3 seeds on claude-haiku, fixture store with keyword-overlapping confusers), the memory retrieval substrate scored precision@5 = 100% (9/9) with 100% poisoned-entry rejection, and the retrieval-on arm passed 27/27 model-scored tasks vs 2/27 with retrieval off and 4/27 with a placebo injection. Known limit stays published: mean tie-set 4.111 means top-k ties break by store order, not relevance (the ADR-116/FTS5 signal). SOLE RECORD for this artefact as of 2026-07-25: a second entry (second-brain-retrieval-precision) described the same measurement from the precision angle and had drifted to 5/27, 5/27 and tie-set 3.3 — figures absent from the shared artefact. Two entries over one artefact is what allowed them to disagree while both resolved, so the pair was folded into this one. |
quant | internal/bench/reports/second-brain-retrieval.json#retrieval-on |
✅ |
| 120 governed rules. | quant | exec:check_artefact_count_messaging -> 0 |
✅ |
Scope de-duplication of the rule projection removes 38.0% of the median cold-start payload (87,677 of 230,556 tokens) against a pre-registered 15% bar. CONDITION, inseparable from the number: that figure is FIXTURE-MEASURED on byte-identical global and project projections, and it is currently unreachable for production installs — the installer stamps ownership metadata (package: / source_path:) onto every installed rule unconditionally, so the two scopes are produced by two writers with deliberately different output, the byte-identity gate correctly refuses to dedup, and the recipient set is empty. The mechanism works; the saving is not realised by any consumer today. |
quant | agents/settings/contexts/cache-economy-refusals.md#Honest null — scope de-duplication is measured but |
✅ |
| On a deterministic multi-session recall corpus, the memory substrate produces a measured, placebo-controlled recall lift — memory-on 27/27 vs no-memory 10/27 and vs equal-byte placebo 9/27 (claude-haiku-4-5, n=9 tasks x 3 seeds, sign test p=0.031 for BOTH pairings). Scoped honestly: this is the context-value upper bound (perfect retrieval on a one-fact-per-task corpus), not retrieval precision under a large store. | quant | internal/bench/reports/second-brain-delta.json |
✅ |
sequential-thinking applies chain-of-thought decomposition with two constraints the original formulation does not carry — a cap on the number of thoughts and a mandatory validation step — specifically to bound the unbounded-expansion failure mode. |
qual | https://arxiv.org/abs/2201.11903 (2026-08-11) |
✅ |
Every artifact the package ships — source AND the condensed projection that reaches consumers — is machine-scanned in CI for hidden-Unicode, mixed-script-confusable, and instruction-smuggling payloads (the rules-file-backdoor class); a finding blocks the release before npm publish, not just the merge. |
qual | exec:lint_agent_security -> 0 |
✅ |
| 299 skills. | quant | exec:check_artefact_count_messaging -> 0 |
✅ |
skill-improvement-pipeline converts post-task outcomes into durable written lessons rather than in-context retries, following the Reflexion formulation. |
qual | https://arxiv.org/abs/2303.11366 (2026-08-11) |
✅ |
Every cross-skill SKILL.md link in the authored skill corpus resolves on disk, and the census behind that statement is derived by the same collector the gate scans with. |
quant | agents/evidence/metrics/skill-link-census.json#"dead_links": [] |
✅ |
Whether serving low-priority skills over MCP instead of listing them natively improves skill selection (H1) is NOT established, and projection.mode: tiered therefore stays opt-in. |
quant | agents/evidence/analysis/skill-tiering-matrix-arm.md#The question this arm cannot answer |
✅ |
On a default Claude Code install projection.mode: tiered costs MORE standing context than legacy-all — 2,259 tokens against 1,956 — because the host already caps description delivery at roughly the Tier A set; the 82% saving exists only against a 100%-delivery counterfactual. |
quant | agents/evidence/analysis/skill-tiering-matrix-arm.md#H2 |
✅ |
| Removes only its own keys from a shared host config (matched by JSON-pointer + SHA-256), never a neighbour tool’s entries. | qual | exec:vitest run tests/lib/json_pointers.test.ts -> 0 |
✅ |
| On the pre-registered 12-fixture defect-finding corpus (three arms, deterministic file-level recall, codex reviewer gpt-5.5), the cross-model team-review arm produced NO recall lift over single-model self-review — all three arms recalled every planted defect (Δ = 0, H1 not met). Honest null, ceiling-limited (recall 1.00 everywhere: the seeded defects are too obvious to discriminate the arms on recall); the only non-null signal is a single self-review false positive on the controversial-but-correct control vs 0 for team/council. No cross-model quality/lift claim binds; team mode stays workflow-value-only. Re-open: a judge-survivable-subtlety corpus or a new model generation. | quant | internal/bench/reports/defect-finding.json#honest_null |
✅ |
A MEASURED-BUT-NOT-SHIPPED experiment — a lean_projection.mode: delivery projection plus the default-OFF rule-inject hook concern delivers rule bodies on a trigger match at 579/579 byte-equal deliveries, 94/94 labelled rules reachable, 0 false fires over 194 labelled near-misses, and $0.7285 vs $4.0335 per 50-turn x 5-spawn session at sonnet rates (standing rule corpus 120,582 -> 18,573 exact-BPE tokens). Shipped default does NOT include this: lean_projection.mode still resolves to eager-all, the concern emits zero bytes, and the flip is Claude-only and carries an unpaid activation charge (a 20,480-byte emission row above the 4,096/2,048-byte slot sums). Latency is not part of that charge: the concern carries no tokenizer and measures p95 0.04-0.05 ms gate-closed and 0.61 ms gate-open, with the whole pre_tool_use slot at 62 ms against a 175 ms budget. What these four endpoints license is delivery equivalence and cost, and nothing wider — the paired-judging instrument is closed by ADR-202 at inter-evaluator Cohen’s kappa 0.472 against a registered 0.800 floor, and this run does not reopen it. Method: model_rule_injection --endpoints over the frozen tests/eval/routing-matrix corpus, pre-registered in internal/bench/thin-inject-PREREG.md before the report artefact existed, with --selftest requiring each endpoint to reject a planted defect first. |
quant | internal/bench/reports/thin-inject-2026-08-23.md#endpoints |
✅ |
On the A3 orchestration eval (deterministic verify, measured token deltas), the production-validator subagent returned the correct verdict on both fixtures — NOT READY with the exact file:line citation on a planted hollow implementation, READY with zero spurious findings on the clean control — while consuming ~45k fewer tokens than the inline-host baseline on each task. Scope: two planted fixtures on a Claude Code host, not a broad hit-rate. |
quant | internal/bench/orchestration/pv-a3-results.md#token_delta_provenance: measured |
✅ |
58 backed claim(s) — all evidence pointers resolve in CI.
How many of those re-derive themselves
Section titled “How many of those re-derive themselves”15 of 58 backed claims carry exec: evidence —
CI re-runs the command and compares its exit code to the claim, so a
stale one turns the build red. The other
43 rest on a pointer: CI checks that the artefact
exists and contains what it should, which cannot distinguish a live
claim from one whose producer nobody has run in months.
That residue is not an oversight, and it is listed rather than rounded away. A claim can only re-derive itself when a deterministic command carries the verdict in its exit code. These cannot, for the reasons given:
| Claim | Why it cannot re-execute |
|---|---|
| The two surfaces the provider-lifecycle contract obliges to agree on an adapter tier — the adapter header and | prose or contract artefact — no exit code carries the verdict |
adversarial-review structures its critique as branching exploration with explicit pruning rather than a sing |
external cite — CI does not fetch the network |
The .augment-plugin/ manifest version is the package version, not an independent plugin-API version, and eve |
prose or contract artefact — no exit code carries the verdict |
analysis-autonomous-mode runs iterative self-critique between steps rather than only at the end, following t |
external cite — CI does not fetch the network |
bug-analyzer verifies each candidate root cause against a concrete trigger before reporting it, following th |
external cite — CI does not fetch the network |
| The release process is documented as an inheritable runbook + succession doc, and the project’s bus-factor (tr | prose or contract artefact — no exit code carries the verdict |
| On the pre-registered 2-arm retrieval benchmark (18 hand-verified code-structure questions across 3 real repos | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
On the first post-fix behaviour-conformance measurement (round 5, /analyze:conformance --limit 30, run 2026- |
benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| A MEASURED-BUT-NOT-SHIPPED experiment — the thin rule projection reduced eager rule load 78,513 → 13,881 GPT-t | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| The first cross-vendor parity pass (5 orchestration-corpus tasks × 2 vendors × 3 repeats, identical prompts vi | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| On a 32-file labelled clean-UI corpus, 18 of the 19 shipped design-slop rules produced zero false positives; t | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| On a weak host (claude-haiku-4-5) the package produces a significant, placebo-controlled discipline lift on sc | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| On the READ-ONLY FAN-OUT slice family, tier-downshifted subagent dispatch (lite/haiku vs session-tier-proxy so | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| The text-layer-only boundary above is machine-enforced, not asserted: a scope-guard test fails if any corpus e | prose or contract artefact — no exit code carries the verdict |
| The retrieval sanitize floor covers the TEXT layer only, by construction — never file or network channels (ima | prose or contract artefact — no exit code carries the verdict |
| The lift-carrying essential cut (kernel + downstream-changes) keeps a significant weak-host discipline lift at | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| A credential-free prescription layer reads content the host’s own web tools cannot fetch at all — Reddit threa | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| Four adjacent governance properties were closed as regression tests rather than phases, on the expectation tha | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| The council aggregation cannot be steered against a refusal. Pre-registered spike S0.1 measured that it WAS cl | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| Pre-registered spike S0.2 measured that both of this package’s fail-closed gates judged the shape of ONE actio | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| Pre-registered spike S0.3 measured that a stated uncertainty, hedge, or provenance marker survives this packag | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| Hook dispatch runs as one precompiled node process with all concerns in-process, and the CI latency gate measu | prose or contract artefact — no exit code carries the verdict |
| On the fixture corpus (n = 20 before/after pairs, 16 length-controlled within ±25%), the humanizer pass remove | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
The injection-scan post_tool_use detector measures 99.00% recall (297/300) and a 0.85% false-positive rate ( |
prose or contract artefact — no exit code carries the verdict |
| A fresh registry install of this package carries zero high/critical npm-audit findings on the runtime dependen | prose or contract artefact — no exit code carries the verdict |
The judge-skill family (judge-bug-hunter, judge-code-quality, judge-security-auditor, judge-synthesis, |
external cite — CI does not fetch the network |
| Word-boundary-anchored keyword matching reduced unintended rule activations by 12.5% (495 → 433) over the 302- | prose or contract artefact — no exit code carries the verdict |
| NONE of the backed ledger claims are machine-re-verifiable today — every evidence pointer is checked for exist | prose or contract artefact — no exit code carries the verdict |
| Registering an MCP server with this package costs standing context on every session, and the kernel server’s 2 | prose or contract artefact — no exit code carries the verdict |
| The whole layer is compiled into host agents with zero runtime daemon. | prose or contract artefact — no exit code carries the verdict |
| On a 12-fixture option-decision corpus (3 arms × 2 providers, blind rubric judge claude-opus-4-8, pre-register | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| We publish our own measured null results and retire or constrain features when the evidence does not support t | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| On the live end-to-end retrieval benchmark (9 tasks × 3 arms × 3 seeds on claude-haiku, fixture store with key | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| Scope de-duplication of the rule projection removes 38.0% of the median cold-start payload (87,677 of 230,556 | prose or contract artefact — no exit code carries the verdict |
| On a deterministic multi-session recall corpus, the memory substrate produces a measured, placebo-controlled r | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
sequential-thinking applies chain-of-thought decomposition with two constraints the original formulation doe |
external cite — CI does not fetch the network |
skill-improvement-pipeline converts post-task outcomes into durable written lessons rather than in-context r |
external cite — CI does not fetch the network |
Every cross-skill SKILL.md link in the authored skill corpus resolves on disk, and the census behind that st |
prose or contract artefact — no exit code carries the verdict |
| Whether serving low-priority skills over MCP instead of listing them natively improves skill selection (H1) is | prose or contract artefact — no exit code carries the verdict |
On a default Claude Code install projection.mode: tiered costs MORE standing context than legacy-all — 2,2 |
prose or contract artefact — no exit code carries the verdict |
| On the pre-registered 12-fixture defect-finding corpus (three arms, deterministic file-level recall, codex rev | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
A MEASURED-BUT-NOT-SHIPPED experiment — a lean_projection.mode: delivery projection plus the default-OFF `ru |
benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
| On the A3 orchestration eval (deterministic verify, measured token deltas), the production-validator subagent | benchmark output — regenerating it needs paid or stochastic model calls no CI job can re-derive |
Artefact counts in public prose (skills, commands, governed rules,
guidelines, personas) are generated from source and CI-drift-checked:
update_counts.ts writes the numbers, check_artefact_count_messaging.ts
fails the build on any count-shaped prose mention that drifts from the
source count — or on two different numbers for the same artefact kind.
We also publish our debt: 23 claim(s) are logged as
unbacked inventory in the ledger — not yet bound, and therefore not
allowed to carry a marker in public prose. Hiding them would be the
opposite of the point.
And our nulls: 6 claim(s) are resolved-null —
measured, the threshold was missed, and the entry is closed rather than
left open forever. A null that stays filed as pending debt is a claim
quietly waiting to be re-argued.
2. We publish honest nulls
Section titled “2. We publish honest nulls”Benchmark results — including the runs where the package changed nothing —
live in docs/benchmark.md. We do not delete a measured null to
make a number look better; the null is the evidence of honesty.
The freshest example: the persona-placebo benchmark
(2026-07-12, 3 arms × 2 providers, blind rubric judge, pre-registered
hypotheses) measured whether famous-figure persona framing improves decision
answers over the bare method text it wraps. It does not (Δ=0.17, p=0.607) —
while provider diversity moved judged quality ~15× more than persona identity.
The measured null closed a planned feature (persona panel-mode) instead of
shipping theater; the full per-cell data is committed at
internal/bench/reports/persona-placebo.json.
Behavioural-eval coverage — the honest baseline. Skill quality is only
as good as its measurement. Today 42 of 299 skills carry a behavioural
evals.json; the highest-traffic / highest-cost tiers (default-surface +
rich + routers) are fully covered (35 of 35), the long tail
(7 of 264) is not. We publish that gap rather than imply
“299 evaluated skills”: coverage is measured per tier
(./scripts-run src/scripts/skill_eval_coverage), CI-ratcheted so it can
only rise, and the priority tiers carry a hard tier floor: every
rich / default-surface / router skill MUST have a behavioural eval or an
explicit exemption-with-reason in internal/evals/tier-floor-exemptions.json
— no silent exemptions (the bright line deliberately replaces a
weighted-coverage score). Authoring long-tail evals is gated on per-case
human ratification (a generated assertion that checks the wrong property is
worse than none), so the number grows deliberately, not overnight.
Non-coding domain soundness — scoped, not proven. The finance /
founder / ops / content profiles sell concrete domain value (DCF,
runway, RICE, incident command, messaging), but the skills are forged on
TS/PHP — “promising, not proven” off those stacks. A disclaimer floor
bounds liability, not correctness: a skill can be format-correct,
disclaimered, and still embed a wrong domain assumption. Today 9 of 20
default-surface domain skills carry a sourced domain-truth fixture
(./scripts-run src/scripts/domain_soundness_status); the rest are labeled
unvalidated and the validated count is CI-ratcheted. The fixtures landed
so far are the deterministic targets (runway, unit-economics, DCF,
forecasting, scenario band/sensitivity), whose answer keys are computed
from cited standard formulas — never the skill’s own output; the
rubric targets (incident command, messaging, fundraising, editorial)
need domain-competent grounding and remain unvalidated, so validation
lands deliberately — the gap is published, never implied away.
Second-brain substrate — measured recall lift, honestly bounded. On a
deterministic multi-session recall corpus, the memory substrate beats a
no-memory baseline AND an equal-byte placebo: memory-on 27/27 vs
no-memory 10/27 vs placebo 9/27 (claude-haiku-4-5, 9 tasks × 3
seeds, sign test p = 0.031 for both pairings) — a real,
placebo-controlled lift (internal/bench/reports/second-brain-delta.json).
Scoped, not oversold: this is the context-value upper bound (perfect
retrieval on a one-fact-per-task corpus), NOT retrieval precision under a
large store; the lift concentrates exactly where memory is the only source
and ties where the prompt self-contains the fact. Boundary vs a human PKM,
and why the Obsidian export stays declined, are in
docs/second-brain-scope.md.
A follow-up removed the perfect-retrieval assumption: against a store of keyword-overlapping confusers the REAL retrieval recalls the needed decision into the top-5 (9/9) and the model disambiguates it — retrieval-on 27/27 vs no-memory 5/27 and vs placebo 5/27 (p=0.008 both). The honest limit: the keyword scorer recalls but does not rank (mean tie-set 3.3), which is the SQLite-FTS5 activation signal (ADR-116) at scale.
3. Known limits (published, witness-tested)
Section titled “3. Known limits (published, witness-tested)”Each limitation below is stated by the skill itself and carries a witness test that reproduces it. If the limitation is ever fixed, its witness goes red — so a “Known limit” can never quietly become false.
| skill | known limit | witness |
|---|---|---|
check-refs |
Only validates references to known-root paths (docs/, skills/, rules/, commands/, contexts/, personas/, …). A relative-path link such as ./sibling.md or ../foo.md is not matched, so a broken relative link is never reported. |
tests/scripts/witness/check_refs_relative_gap.test.ts |
4. What is checkable — us vs. the category
Section titled “4. What is checkable — us vs. the category”This is not a takedown — the point is the last column. For each claim,
our evidence is a pointer you can resolve on a fresh checkout; the wider
category is described only by what is publicly observable, never a named
competitor and never a counter-claim to anyone’s headline number. A claim
is “checkable” only when its our evidence pointer resolves — CI enforces
that (task check-comparison), so this column can never lie.
| Claim | Our evidence | The category | Checkable? |
|---|---|---|---|
| No resident runtime — no background daemon, no state database or service, no auto-write memory; deterministic, rebuildable file indexes only. | docs/contracts/no-runtime-boundary.md#file-first |
Swarm-runtime tools in this category ship a background process and/or a resident state store by design; that is an architectural fact of the runtime approach, not a defect. | ✅ |
| Capability/benchmark results are published — including the runs where the package changed nothing. | docs/benchmark.md |
A category headline figure (e.g. an ‘84.8%’ score) appears across marketing surfaces with no reproducible methodology published — so it cannot be verified either way. | ✅ |
| Every public claim binds to machine-checked evidence. | docs/CLAIMS.md |
Marketing claims in this category are not bound to a machine-checked ledger a reader can reproduce. | ✅ |
| Skills publish their own known limits, each with a witness test. | tests/scripts/witness/check_refs_relative_gap.test.ts |
Published known-limitation surfaces backed by reproducing tests are not a standard artifact in this category. | ✅ |
| A 30-second single-file wedge exists — one read-only subagent, one command, verdict with file:line — and its promise is a scoped, published eval, not a feature list. | internal/bench/orchestration/pv-a3-results.md#token_delta_provenance: measured |
Category entry points typically install the full platform (profiles, packs, runtime) before the first felt win; single-artifact entries with a published eval behind the promise are not a standard artifact. | ✅ |
| The persona-identity question was tested and published as an honest null (identity swap ~zero effect, provider choice ~15x larger) instead of being shipped as theater. | internal/bench/reports/persona-placebo.json#honest-null |
Named-persona prompts are shipped as a feature across the category without a published A/B separating identity from method. | ✅ |
| Every upstream-tool install prescription the package ships is version-pinned, intake-recorded, and enforced offline in CI — an unpinned version, a branch/archive install source, or a pipe-remote-to-shell instruction fails the build. | src/scripts/validate_reach_prescriptions.ts |
Capability layers in this category commonly install themselves and their upstream dependencies from a moving target (a branch archive or an unpinned package) via a piped remote installer; whether any given one pins is observable per repo, but a CI gate that fails the build on an unpinned prescription is not a standard artifact. | ✅ |
| Platform reads that the host’s own web tools cannot perform at all (a refused domain, an HTTP 402) are shipped as credential-free prescriptions whose reliability was measured per channel against a native control, with the thresholds frozen before the run and the narrowing caveat published alongside the pass. | docs/benchmark.md#ship-gated-reach |
Reaching gated platforms is commonly solved in this category by asking for an API key or a session cookie, and the capability is asserted from a feature list rather than a per-channel reliability measurement with a native-arm control; publishing the case where the native tools already suffice, next to the pass, is not a standard artifact. | ✅ |
| A code-intelligence engine was built natively, measured against plain disciplined grep, and then permanently defaulted off when it lost — the retrieval null (recall 0.365 vs grep 0.797 on graph-shaped questions) is published, and a deprecation date is on it. | internal/bench/reports/code-graph-vs-grep.md#Verdict — NULL, decisively |
Code-graph / semantic-index retrieval is shipped as a headline capability across this category; a published measurement against a plain-grep control — let alone one that retires the feature the vendor built — is not a standard artifact. | ✅ |
What each one prevents
Section titled “What each one prevents”The same rows, read as failure modes rather than as comparisons. Each names something that goes wrong without the control and carries the identical resolvable pointer — a projection of the table above, not a second list to keep in sync.
| Without it | The control | Evidence |
|---|---|---|
| A background process keeps running after you stop using it, holds state you did not write, and has to be kept alive and upgraded. | No resident runtime — no background daemon, no state database or service, no auto-write memory; deterministic, rebuildable file indexes only. | docs/contracts/no-runtime-boundary.md#file-first |
| You adopt on a headline number nobody can re-run, and never learn which of its measurements changed nothing. | Capability/benchmark results are published — including the runs where the package changed nothing. | docs/benchmark.md |
| A shipped number drifts from the artefact it came from and stays published, because nothing re-checks the claim against its source. | Every public claim binds to machine-checked evidence. | docs/CLAIMS.md |
| A limitation is discovered by you, in your repo, at the moment it costs you something — instead of being stated up front and held by a test. | Skills publish their own known limits, each with a witness test. | tests/scripts/witness/check_refs_relative_gap.test.ts |
| Evaluating the tool means installing the whole platform first, so the first thing you learn is the setup cost rather than whether it works. | A 30-second single-file wedge exists — one read-only subagent, one command, verdict with file:line — and its promise is a scoped, published eval, not a feature list. | internal/bench/orchestration/pv-a3-results.md#token_delta_provenance: measured |
| You pay tokens for prompt theater that was never measured against the method text it wraps. | The persona-identity question was tested and published as an honest null (identity swap ~zero effect, provider choice ~15x larger) instead of being shipped as theater. | internal/bench/reports/persona-placebo.json#honest-null |
| An install instruction silently follows a moving target — an unpinned package or a piped remote script — so what lands in your repo is whatever the source held that day. | Every upstream-tool install prescription the package ships is version-pinned, intake-recorded, and enforced offline in CI — an unpinned version, a branch/archive install source, or a pipe-remote-to-shell instruction fails the build. | src/scripts/validate_reach_prescriptions.ts |
| You hand over a credential — or your account — for a read the tools could already do, and you cannot tell which channels actually work because the capability was claimed rather than measured. | Platform reads that the host’s own web tools cannot perform at all (a refused domain, an HTTP 402) are shipped as credential-free prescriptions whose reliability was measured per channel against a native control, with the thresholds frozen before the run and the narrowing caveat published alongside the pass. | docs/benchmark.md#ship-gated-reach |
| You adopt an indexing engine on the promise that structure beats search, and nobody ever measured it against the search it replaces — so the index keeps costing build time and staleness risk for retrieval that got worse. | A code-intelligence engine was built natively, measured against plain disciplined grep, and then permanently defaulted off when it lost — the retrieval null (recall 0.365 vs grep 0.797 on graph-shaped questions) is published, and a deprecation date is on it. | internal/bench/reports/code-graph-vs-grep.md#Verdict — NULL, decisively |
4b. The two existing axes — enforcement level per rule, evidence form per claim
Section titled “4b. The two existing axes — enforcement level per rule, evidence form per claim”Pure projection of what the repo already knows — the enforced_by
resolution (check_enforcement_coverage) and the claims ledger
(docs/CLAIMS.md). No new taxonomy, zero hand-written rows.
Axis 1 — enforcement level per rule. 120 rules · 15 blocking (12.5%) · 10 observer · 0 local-only · 82 undeclared (no enforced_by yet).
denominator: 120 rule(s), frame in-scope (src/rules/*.md) == governed-total 120
| Rule | Effective level | Declared backstop(s) |
|---|---|---|
active-remediation |
none | instruction-only: the note, the ask and the user decision are all prose, so no gate can tell a discharged issue from a mentioned one |
code-provenance |
none | instruction-only: close-the-source-and-re-derive is a pre-write reasoning step only the model observes; CI checks the ledger, never the derivation |
context-hygiene |
observer | hook:context-hygiene |
council-availability |
none | instruction-only: no gate reads a chat claim about availability; check_council_config_location covers the tree side only |
decision-revisit-gate |
none | instruction-only: no gate can observe an agent citing a decision it never opened; adr_cite_check is deterministic where it runs and nothing makes it run |
design-review-after-ui-write |
none | instruction-only: no artefact proves a design review happened outside the work-engine dispatcher; the review verdict is self-report |
evaluator-independence |
observer | hook:evidence-independence |
fix-what-you-see |
none | instruction-only: ownership-as-excuse is a disposition in prose; no gate can see a red check handed back with its cause named |
framework-neutrality-in-generic-skills |
validator | validator:src/scripts/lint_framework_leakage.ts |
git-history-discipline |
hook | hook:block-no-verify |
language-and-tone |
validator | validator:src/scripts/check_md_language.ts |
lethal-trifecta-guard |
validator | validator:src/scripts/lint_skill_frontmatter_safety.ts |
media-governance-routing |
validator | validator:src/scripts/lint_media_policy_linkage.ts |
minimal-safe-diff |
observer | hook:minimal-safe-diff |
missing-skill-recovery |
none | instruction-only: nothing can observe an agent concluding that no skill exists; the skill-route concern covers only the prompts where the ranker is confident |
no-roadmap-references |
validator | validator:src/scripts/check_no_roadmap_refs.tsvalidator:src/scripts/check_council_references.ts |
non-destructive-by-default |
none | none |
onboarding-gate |
observer | hook:onboarding-gate |
output-discipline |
validator | validator:src/scripts/lint_output_slop.ts |
persona-governance |
validator | validator:src/scripts/lint_persona_governance.ts |
playbook-precedence |
none | instruction-only: no gate can tell a playbook-first run from a skill-first one — both produce a diff, and which answer was consulted leaves no artefact |
preservation-guard |
validator | validator:src/scripts/check_condensation.tsvalidator:src/scripts/skill_linter.ts |
recurring-criticism |
none | instruction-only: the earlier disposition, the three outcomes and the hardening are all prose; the self-repair occurrence counter covers only detector-matched defects |
roadmap-progress-sync |
observer | hook:roadmap-progress |
secret-vcs-guard |
validator | validator:src/scripts/check_secret_leak.ts |
security-sensitive-stop |
none | instruction-only: threat-model-before-you-edit is a pre-edit reasoning step only the model observes |
self-repair-loop |
observer | hook:self-repair |
session-canary |
observer | hook:session-canary |
settings-ask-protocol |
none | instruction-only: no gate counts the questions in a chat turn; settings:set fences the illegal WRITE, never the badly-shaped ask |
skill-quality |
validator | validator:src/scripts/skill_linter.ts |
source-confidentiality |
validator | validator:src/scripts/check_no_external_sources.ts |
source-of-truth |
validator | validator:src/scripts/check_condensation.ts |
telegraph-speak |
observer | observer:maintainer-review |
token-optimizer-maintenance |
validator | validator:src/scripts/check_token_optimizer_freshness.ts |
tool-safety |
validator | validator:src/scripts/lint_agent_security.ts |
ui-audit-gate |
observer | hook:design-pass |
untrusted-input-defense |
none | instruction-only: no deterministic gate inspects fetched content for injected instructions; injection_scan_hook is warn-only and default-OFF |
verify-before-complete |
observer | hook:verify-before-complete |
Undeclared rules (82) carry no row — an honest gap beats a false claim.
Axis 2 — evidence form per public claim. 87 ledger entries · 58 backed · 23 unbacked inventory · 6 resolved-null.
| Claim id | Kind | Status | Evidence pointer |
|---|---|---|---|
adapter-lifecycle-day-one-table |
qual | backed | docs/contracts/provider-lifecycle.md#Current tier assignment |
adr-beneficiary-grade-bias |
quant | unbacked | PRE-REGISTERED 2026-08-21 (road-to-evidence-based-adr-governance Phase 6.2). WHY THIS EXISTS: it is the measured form of the conflict of interest a two-seat council named as the design's central risk. An agent grading a record that constrains it has a structural incentive to grade it weak, and recording provenance documents the conflict without neutralising it — an audit log is not a control. METRIC: partition accepted records into CONSTRAINING (the record forbids, gates, or requires human involvement in an action an agent could otherwise take) and NON-CONSTRAINING, classified from the record's Decision section by a party that did not assign its grade; compare the E0/E1 rate across the two partitions. THRESHOLD, fixed now: the constraining partition's weak-grade rate may exceed the non-constraining partition's by no more than 15 percentage points. POWER: at least 25 records per partition, else UNDERPOWERED. FALSIFICATION: (1) a breach is evidence of grade deflation and suspends any authority consequence outright — it does not trigger a re-grade by the same party, which would be the same conflict applied to its own audit; (2) the reverse skew (constraining records graded STRONGER) is reported as-is and is not read as a virtue: it would suggest defensive over-grading, which inflates reopen burden instead of deflating it; (3) a null result here does NOT license the authority coupling on its own — it removes one objection out of four (claim:adr-grade-accuracy-vs-gold, claim:adr-evidence-discovery-recallandclaim:adr-interruption-baseline are the others), and the coupling stays owner-reserved regardless. |
adr-evidence-discovery-recall |
quant | unbacked | PRE-REGISTERED 2026-08-21 (road-to-evidence-based-adr-governance Phase 6.2). WHAT IT MEASURES: the failure mode the discoveryfield exists for. A bare E0 collapses five states — evidence absent, evidence existed and was never cited, cited somewhere non-standard, present in the tree and not found, external and never fetched — and the last four are discovery failures, not evidence failures. A record graded weak because nobody looked is the cheapest possible way to manufacture a reopenable lock. METRIC: draw a random sample of at least 15 records carryingstrength: E0withdiscovery: complete, run a deeper independent search on each (full-tree grep for the decision's own terms, the roadmap and PR that produced it, the external sources its body names), and count how many turn out to have locatable evidence. THRESHOLD, fixed now: no more than 20% of the sample may turn out to have findable evidence. FALSIFICATION: (1) a breach means completeis being asserted whereincompleteis the honest value, and the remedy is to makeincompletethe only permitted value for a heuristic proposal rather than to lower the floor; (2) if nearly every honest answer isincomplete, that is itself the finding — the field then records uncertainty rather than resolving it, which is stated in ADR-240 § Assumptions as an accepted possibility rather than discovered later. HONEST-NULL PATH: if no record ever carries discovery: complete, the recorded finding is that the strong form was never used, and this claim resolves null with the observation that E0 in this tree means "unsearched" in practice. |
adr-grade-accuracy-vs-gold |
quant | unbacked | PRE-REGISTERED 2026-08-21 (road-to-evidence-based-adr-governance Phase 6.2). WHY NOT INTER-REVIEWER AGREEMENT, which is the obvious metric and the wrong one: agreement measures shared bias as readily as accuracy, and this repository has the receipt — 44 engine-shaped REJECT records accumulated under correlated council agreement (engine-reclassification-2026-07.md) and were disposed of by one measurement (claim:code-graph-retrieval-null). Two reviewers who searched the same way and read the same rubric will agree while both being wrong. METRIC: an anchor sample of 12-15 records is graded independently, then adjudicated to a gold value by a party that did not produce either grade; accuracy is the proportion of proposals matching gold on the E0/E1-versus-E2+ boundary, which is the boundary that matters because it is the one the burden table prices. THRESHOLD, fixed now: 85%. Reported WITH the disagreement count and stratified by record type, never as a bare percentage. POWER: fewer than 12 adjudicated records is UNDERPOWERED and yields no claim. FALSIFICATION: (1) high accuracy on records that constrain nothing, paired with low accuracy on records that constrain agent behaviour, is a FAILURE even if the aggregate clears 85% — see claim:adr-beneficiary-grade-bias; (2) an adjudicator who saw a proposal first is not independent and that sample is void; (3) if adjudication itself proves unrepeatable, the finding is that evidence grading is not reliably gradeable here, which is a publishable null and closes the authority question by itself. |
adr-interruption-baseline |
quant | unbacked | PRE-REGISTERED 2026-08-21 (road-to-evidence-based-adr-governance Phase 6.1). REGISTERED BEFORE THE MECHANISM CHANGES ANY BEHAVIOUR, and that ordering is the point: the axes ship descriptive (ADR-240 § 2 — a grade prices review burden and confers no authority), so the baseline is measured against a tree where nothing has been unlocked yet. METRIC, fixed now: a contact counts when a run stops or asks the owner AND the stated reason names an ADR — read from session transcripts, not from an agent's self-report, because "I was blocked by ADR-X" is exactly the claim the surrounding roadmap exists to stop taking on faith. DENOMINATOR: 20 roadmap runs, counted as /roadmap:process-* invocations that reach at least one closed step; a run that aborts before its first step is not a run. POWER: at least 20 runs on each side, else UNDERPOWERED and no claim in either direction — a five-run sample producing a favourable ratio is the failure this line exists against. FALSIFICATION: (1) a fall in contacts accompanied by a rise in the held-defect rate is NOT a win and is reported as a trade, not a success; (2) the comparison is rate-vs-rate, never absolute counts, since run volume will differ; (3) a rise is published as a rise — this metric can close the mechanism, and per ADR-240's own review_trigger a flat result reopens that record. HONEST-NULL PATH: if 20 runs do not accumulate, the recorded finding is that the metric was pre-registered and never had a population, and the interruption claim is withdrawn rather than left pending. |
adversarial-council-finding-coverage |
quant | resolved-null | docs/benchmark.md#adversarial-verification-council |
adversarial-review-tree-of-thoughts |
qual | backed | https://arxiv.org/abs/2305.10601 (2026-08-11) |
augment-manifest-version-package-synced |
qual | backed | src/scripts/lint_marketplace.ts#check_augment_manifests |
autonomous-analysis-self-refine |
qual | backed | https://arxiv.org/abs/2303.17651 (2026-08-11) |
budget-routing-relation |
qual | resolved-null | docs/contracts/budget-routing.md#Why it was retired |
bug-analyzer-chain-of-verification |
qual | backed | https://arxiv.org/abs/2309.11495 (2026-08-11) |
bus-factor-tracked |
qual | backed | docs/succession.md#trailing 90 days |
code-graph-retrieval-null |
quant | backed | internal/bench/reports/code-graph-vs-grep.md#Verdict — NULL, decisively |
command-count |
quant | backed | exec:check_artefact_count_messaging -> 0 |
conformance-advisory-vs-blocking |
quant | backed | src/domains/analysis-workbench/analyze/conformance/command.md#Both blocking carriers reached zero |
context-fidelity-compaction-compliance |
quant | unbacked | PRE-REGISTERED 2026-08-17 (road-to-context-fidelity Phase 0 Step 4 — registered while the census that would answer it, cf01, has NOT been run, so this is a genuine pre-registration and not a ledger entry written around a number already in hand). THRESHOLD fixed before data, taken verbatim from the roadmap: **a baseline compliance at or above 90 % for all three probe classes closes Phase 1 UNBUILT** and the null is published with the host version recorded. The three probe classes are a session-canary-bound obligation, a completion-gate reminder, and one trigger-loaded rule with a detectable obligation; the threshold binds all three separately, so a mean of 90 % carried by one strong class is not a pass. METHOD CONSTRAINT discovered before the census ran, and it is why this entry does not simply wait: cf03 measured 29 compaction events across 473 sessions, **all 29 tagged host-automatic and none tagged manual** (agents/evidence/eval-findings/context-fidelity-cf03.md). That zero is **absence of a RECORD, not absence of an event** — corrected on R2 finding 6, which caught the first phrasing here reporting an unobservable as an observation. The detector is pinned to one OBSERVED auto event (src/scripts/_lib/session_eol.ts:11-19) and nothing in the tree establishes that a manual compaction writes a compact_boundary record at all. So the constraint on cf01 is sharper than "measures a rare path": until manual detectability is established, a cf01 null is UNINTERPRETABLE — indistinguishable from a compaction that happened and left no trace. Establishing it is one manual compaction in one instrumented session, and it is a precondition rather than a result. FALSIFICATION fixed before data: (1) the host version is stamped on every observation, because compaction survival is a host fact that changes without notice; (2) a probe present only as paraphrase counts as NOT followed — the obligation is the behaviour, not the recall; (3) at or above 90 % on all three classes the recorded consequence is that Phase 1 is not built and the folklore is named as folklore, with the same force as the positive direction. HONEST-NULL PATH is therefore the DEFAULT outcome of a high reading, not a fallback: this claim exists to be able to close work rather than to justify it. |
context-fidelity-memory-staleness |
quant | unbacked | PRE-REGISTERED 2026-08-17 (road-to-context-fidelity Phase 0 Step 4). THE THRESHOLD PREDATES THE DATA, THE LEDGER ENTRY DOES NOT, and the distinction is stated rather than blurred: the roadmap fixed **a stale ratio below 10 % shrinks Phase 2 to stamps only** on the day it was written, BEFORE any census ran; this entry was written after cf02 produced a first reading, so it records a resolution-in-progress rather than claiming the ledger row itself came first. FIRST READING, from agents/evidence/eval-findings/context-fidelity-cf02.md: 107 curated entries walked one by one against the tree — 73 still-true, 23 stale, 11 unverifiable, i.e. **21.5 % of all entries and 24.0 % of the verifiable subset**. Both denominators clear 10 %, so the kill criterion does not fire on either reading and the ladder stays justified. THE LOAD-BEARING FINDING is that the shipped instrument disagrees: memory_reportreportsstaleness-rate=0.0%because all 107 entries carry the SAMElast_validated: 2026-07-09and the SAMEreview_after_days: 365 — one bulk stamping event, so the age axis cannot read stale before 2027-07-09. Reading the kill criterion off that 0.0 % would have closed Phase 2 on a number that measures stamping rather than truth, which is the already-satisfied-test failure this repository has recorded before. WHY THIS STAYS UNBACKED despite having a number: the tree axis was walked BY HAND because no store-wide contradiction sweep exists (check_memory_contradictiontakes–type –key –body, i.e. it validates one proposed entry), three observers each classified one store, and inter-rater agreement is therefore UNMEASURED. A hand classification is not a reproducible instrument, and a ratio that cannot be re-derived by a command is not backing. BACKING TRIGGER: a store-wide sweep exists AND reproduces a ratio within its own stated error of 21.5 %. FALSIFICATION fixed before data: (1) unverifiable entries are counted as their own class and folded into neither side — folding them into still-true inflates the pass rate, folding them into stale manufactures defects; (2) the commit anchor is the precondition for any automated reading, because without it a date cannot be tied to a tree state; (3) below 10 % on a reproducible sweep the recorded consequence is the roadmap's own: the ladder is unbuilt, only stamps ship, and the null is published. |
context-token-reduction |
quant | backed | internal/bench/reports/token-baseline.json#eager_rule_load |
council-fallback-loses-zero-seats |
qual | backed | exec:vitest run tests/scripts/ai_council/council_cli.test.ts -> 0 |
council-parse-outcome-corpus-rate |
quant | backed | exec:vitest run tests/scripts/ai_council/parse_corpus.test.ts -> 0 |
council-vs-solo-baseline |
comparative | unbacked | PRE-REGISTERED 2026-07-12 (road-to-feedback-8.11 Phase 3 — no goalpost-moving after the numbers land; design at docs/design/council-vs-solo-baseline.md). Falsification criteria fixed BEFORE data: (1) quality = blind post-hoc grading against known ground-truth dispositions, two blind judges, admissible only at Cohen's κ ≥ 0.60 (reuse check_quality_regression.ts kappa machinery); (2) the five feedback-proposed admission dimensions are recorded per decision AT pre-registration, so "≥2-of-5" is a testable post-hoc correlate, never a pre-imposed gate; (3) NO lift on any subset (overall, per impact class, per dimension stratum) → honest null, deliberation-protocol phases stop (maintenance-only), recorded in road-to-opt-council-deliberation; lift on a subset → admission criteria derived FROM that subset's characteristics. Execution is spend-gated (user confirms rendered estimate in-session); shadow-log was absent/empty at design time — zero prior council-vs-solo data exists. |
critic-protocol-load-bearing-ab |
quant | resolved-null | internal/bench/adversarial-council/runs/critic-protocol-ab-report.json |
cross-model-parity-count |
quant | backed | internal/bench/reports/parity-count.json#min over hosts of median |
cross-source-consistency-precision |
quant | unbacked | PRE-REGISTERED 2026-07-28 (road-to-feedback-9.2.0-followups Phase 1 — no goalpost-moving after the numbers land). Falsification criteria fixed BEFORE data: (1) fixtures + expected actions are pinned in internal/bench/corpora/honesty-false-premise.yaml(shared with the honesty bench, extended-not-forked); (2) the scorer issrc/scripts/bench_cross_source_eval.ts (ask|proceed|warn classification, forbidden-assumption + over-firing checks) — precision = correctly-surfaced discrepancies / all surfaced; over-firing = asks on negative controls / negative controls; (3) the run needs real model responses per fixture (paid, maintainer-gated spend) — this entry stands as documented debt until that run lands; (4) HONEST NULL consequence bound: precision < 85% or over-firing > 5% → loosen the rule's default (consistency.cross_source: on→auto) or tighten its confidence tiers — never silently keep firing. This binds the weaker-evidenced default-on rule to a measurement like every other default-flip. |
default-install-context-cost |
quant | backed | exec:update_counts --check -> 0 |
design-slop-false-positive-baseline |
quant | backed | internal/bench/corpora/design-slop-fp-PREREG.md#The ceiling, declared before the run |
discipline-lift-weak-host |
quant | backed | docs/benchmark.md#weak-host-specific |
domain-soundness-scoped |
qual | backed | exec:domain_soundness_status --check -> 0 |
domain-soundness-validated-count |
quant | backed | exec:domain_soundness_status -> 0 |
downshift-cost-reduction |
quant | backed | internal/bench/routing-downshift/results-2026-07-08.md#FAMILY-SCOPED PROVE |
encoding-corpus-scope-guard |
qual | backed | tests/scripts/encoding_corpus.test.ts#FAILS when a deliberately out-of-scope fixture is added |
encoding-floor-text-layer-only |
quant | backed | agents/evidence/reports/encoding-floor-measurement.md#Selected branch: ADOPT |
enforcement-coverage-resolved |
quant | backed | exec:check_enforcement_coverage --check -> 0 |
enforcement-undeclared-denominator |
quant | backed | exec:check_enforcement_denominator -> 0 |
essential-tier-cost-factor |
quant | backed | docs/benchmark.md#REPLICATION FAILED |
eval-coverage-ratcheted |
qual | backed | exec:skill_eval_coverage --check -> 0 |
experiment-loop-iteration-floor |
quant | unbacked | PRE-REGISTERED 2026-08-17 (road-to-metric-loop-and-review-integrity Phase 0/5 — the floor was fixed before the spike ran, and the spike's kill criterion was its complement: "fewer than five clean iterations" would have left Phase 3 unbuilt). Measured on a toy metric in a scratch repository, agents/evidence/eval-findings/metric-loop-s01.md: 6 clean iterations, metric 24 → 3, with iteration 5 reverting a change that improved the metric 67 % and broke behaviour. That result answers the PHASE GATE and is why the skill shipped. It does NOT back this claim, and the distinction is the whole reason the entry stays unbacked: the run was one agent, one session, one toy metric whose evaluator was written alongside the loop, so it measured whether the PROTOCOL holds, not whether the shipped skill drives a real metric. BACKING REQUIRES: ≥ 3 runs of the shipped experiment-loop skill against metrics that existed before the run, each with its register committed, each reaching ≥ 5 clean iterations. DROP: any run below the floor publishes the null and the skill is withdrawn rather than the floor lowered — lowering a pre-registered floor after seeing the data is the tuning this roadmap's own s04 finding forbids. |
forensics-pack-value |
quant | unbacked | agents/evidence/release-findings/ |
gated-platform-reads |
quant | backed | docs/benchmark.md#ship-gated-reach |
governance-adjacent-properties |
quant | backed | internal/bench/reports/governance-invariants.json#adjacent_properties |
governance-aggregation-refusal-invariance |
quant | backed | internal/bench/reports/governance-invariants.json#s0_1_aggregation_steerability |
governance-decomposition-effect-boundary |
quant | backed | internal/bench/reports/governance-invariants.json#s0_2_decomposition_laundering |
governance-marker-preservation-null |
quant | backed | internal/bench/reports/governance-invariants.json#s0_3_marker_survival |
hook-dispatch-latency |
quant | backed | docs/hook-latency.json#invocation_path |
host-agent-count |
quant | backed | exec:vitest run tests/install/toolDetection.test.ts -> 0 |
humanizer-tell-reduction |
quant | backed | internal/bench/reports/humanizer-v1.md#prefers the humanized text |
injection-scan-corpus-rates |
quant | backed | agents/evidence/reports/injection-detector-wiring.md#The numbers |
install-audit-clean |
quant | backed | .github/workflows/release-validation.yml#npm audit --omit=dev --audit-level=high |
judge-family-llm-as-a-judge-foundation |
qual | backed | https://arxiv.org/abs/2306.05685 (2026-08-11) |
keyword-anchoring-census |
quant | backed | agents/evidence/analysis/anchoring-independent-replay-2026-08.md |
lean-init-cost-reduction |
quant | unbacked | PRE-REGISTERED 2026-07-28 (road-to-lean-agent-init Phase 3 — registered BEFORE any savings number is cited anywhere; family-scoped, modeled on downshift-cost-reduction; quality definition reused from the correctness-comparison acceptance, no second truth). Falsification criteria fixed BEFORE data: (1) correctness floor — primitive answer ≡ agent answer on the golden corpus (internal/bench/lean-init/results-2026-07-28.md, 12/12); ANY mismatch on a routed real task recorded via correctness_match: false counts against the claim; (2) negative control — a non-lookup task never routes to a primitive (LOOKUP_CORPUSlk-n1..n4, FP=0); (3) cost metric — read fromagents/runtime/state/audit/*.jsonlorchestration lines taggedorigin: lean-init-2026withlookup_class != null, comparing route_taken: primitivetoken cost againstroute_taken: subagentlines of the same class (n and family scope stated at backing time); (4) segregation — lines carryorigin: lean-init-2026so theroad-to-orchestration-scope-decision sample stays uncontaminated (council Q5, 2026-07-28). PROVE → flip to backed for the lookup family only; DROP → honest null, primitives stay (correctness-validated) but no savings number is ever cited. |
ledger-exec-verifiability |
quant | backed | internal/reports/exec-evidence-feasibility.json#"exec_feasible" |
lexical-ranking-lift |
quant | backed | exec:measure_lexical_ranking -> 0 |
mcp-registered-server-standing-cost |
quant | backed | agents/evidence/metrics/mcp-tool-standing-cost.jsonl#tool_search_threshold |
no-runtime-daemon |
qual | backed | docs/contracts/no-runtime-boundary.md#file-first, no-runtime suite |
orchestration-dispatch-net-win |
comparative | unbacked | PRE-REGISTERED 2026-07-11 (road-to-orchestration-scope-decision Phase 1 — no goalpost-moving after the numbers land). Falsification criteria fixed BEFORE data: (1) held quality is deterministic, scored by src/scripts/check_quality_regression.tsthresholds — a token/wall win that degrades output below the regression threshold FAILS the claim; (2) negative control —pv-02-negative-controlmust NOT trigger dispatch (a classifier that fires on everything is a cost leak, not a win); (3) win metric — ≥15% reduction in token-or-wall onorch-02+orch-03vs the single-agent baseline, read fromagents/runtime/state/audit/*.jsonlorchestration lines throughgateVerdict()/resolveShippedDefault(). Binds to a resolving report once ≥20 real ask-mode telemetry lines exist (Phase 2 — maintainer-run; the corpus –runagent-spawn is gated out of auto-mode). PROVE → flip to backed for the proven family only; DROP → renewed honest-null, keepask, demote orchestration from the public value proposition. |
orchestration-observed-dispatch-cost |
comparative | resolved-null | internal/bench/orchestration/backfill-2026-08-07-verdict.md#honest null |
persona-identity-placebo-null |
quant | backed | internal/bench/reports/persona-placebo.json#honest-null |
plan-gates-measurement-protocol |
quant | unbacked | docs/contracts/plan-review-gates.md#Advisory window (Stage A, verdict #20) |
positioning-honest-nulls |
qual | backed | docs/benchmark.md#honest |
provenance-detector-transformation-sensitivity |
quant | resolved-null | PRE-REGISTERED 2026-07-28 BEFORE the S0.3 baseline run (road-to-provenance-and-license-governance Phase 0; corpus frozen at content-sha256 dbbc84a7325e4fa38483ba05d35d9c0c98fa822ae25d873bd5efbafaf2534bb3 over internal/bench/provenance/, 36 files). Thresholds fixed BEFORE data, per the roadmap's S0.2 and its denominator fix: (1) detector recall on the verbatim+rename-only subset >= 10/16 (8 verbatim + 8 rename-only); (2) false positives on the 12 independent controls <= 1/12; (3) rename-only samples MUST hit (principle 6 — laundering by rename cannot clear a hit); (4) structural-rewrite samples form the residual class and their recall feeds the Phase-5 drop gate (>= 21/24 on the full seeded corpus DROPS Phase 5). The floor is a GO/NO-GO gate for building the CI layer, never the marketed capability — the marketed capability is the measured rate published per S3.1/S3.3 with the scope bound above co-located. HONEST-NULL consequence (K1): thresholds missed => no deterministic-gate claim ever, the behavioural layer ships alone, null published; no silent threshold adjustment. |
provenance-gate-effectiveness |
qual | unbacked | PRE-REGISTERED 2026-07-28 (road-to-provenance-and-license-governance Phase 3, S3.1 — registered AFTER Gate G0's verdict, so this claim's text already reflects the re-scope rather than describing a capability that was later cancelled; the original S3.1 draft text ("AC's provenance gate detects seeded verbatim and rename-only OSS copies at the Phase-0 measured rate") is FALSE post-G0 and is not reused). G0 honest-null context (see provenance-detector-transformation-sensitivityfor the pre-registered thresholds): the deterministic scan layer (jscpd offline + SCANOSS online) measured, on the frozen synthetic corpus, verbatim+rename-only recall 12/16 (union) and false positives 2/12 (union) — missing BOTH the recall and FP thresholds — with SCANOSS alone recalling rename-only samples 0/8. Council decision 2026-07-28 (K1 literal, Option A): nolint_code_provenance.tsships in ANY form in CI, not even advisory — the scan capability exists ONLY as thelicense-compliance-auditskill (src/skills/license-compliance-audit/), invoked deliberately by a human, never wired into any pipeline. Falsification criteria fixed at registration: (1) any deny-class or unknown-license ledger entry passesci(alint_provenance.tsregression); (2) any ledger entry missing atransformation_notepassesci; (3) any rename-only-phrased transformation_note(the 15-phrase rejection list) passeslint_provenance.ts; (4) any user-facing surface asserts or implies a CI-facing similarity/duplication detector exists, or omits the co-located scope statement (S3.2/S3.3 global consequence bound, enforced by lint_provenance_vocabulary.ts). Backing trigger: (a) >=1 real (non-empty, non-fixture) ledger entry survives lint_provenancein a merged PR with a substantive transformation_note, AND (b) the Phase-4 dogfood self-audit (S4.1) publishes its findings against this repo itself, whatever the outcome. Part (b) is already satisfied —internal/bench/provenance/reports/self-audit-2026-07-28.mdpublished a headline finding that 551 of 552 online-scanner hits on this repo's OWN source were self-matches against its own published releases, a third independent argument for the G0 verdict and evidence the corpus's 2/12 false-positive rate understated the real-world surface. Part (a) is NOT yet satisfied —provenance/borrows.jsonl holds zero entries — so this claim stays unbacked until a real borrow lands and clears the ledger. |
reference-loop-upgrade-value |
quant | unbacked | PRE-REGISTERED 2026-08-12 (road-to-cross-repo-differential-loop Phase 6 — no goalpost-moving after the runs land). Falsification criteria fixed BEFORE data: (1) the comparison is against a shadow run of the pre-upgrade command text on the same reference, so "could not have produced" is decided by diffing two documents, not by assertion; (2) a finding counts only with a concrete file:line on OUR side — a probe recording consumer not locatableis an honest result but not a positive; (3) TIME BOUND: 180 days from merge — an event-bound measurement on a rare event is an unbacked row that never settles, so window expiry counts as the bar not cleared; (4) HONEST NULL consequence bound, asymmetric by construction: bar not cleared or window expired → the interop-probe, convergence and–deep mechanisms revert and the null is published; the anchor-table and bound-claim-gate mechanisms stay regardless, because they enforce ADR-211 C/D and the claims ledger — doctrine that already binds — rather than claiming new value. A measurement that never arrives therefore cannot leave the command worse than before the upgrade. |
retrieval-substrate-live-pass |
quant | backed | internal/bench/reports/second-brain-retrieval.json#retrieval-on |
review-independence-changes-consumption |
qual | unbacked | PRE-REGISTERED 2026-08-17 (road-to-metric-loop-and-review-integrity Phase 2/5 — registered BEFORE any consumption claim is made anywhere). The MECHANISM shipped and is machine-checked: check_review_schemarefuses an artifact whoseacceptance_statuscontradicts itsreview_independence, and its –self-testplants the exact defect (a same-family set claimingaccepted) and confirms the rejection fires. The EFFECT is a different question and is not measured: nothing yet observes a consumer reading the field and behaving differently, and the only committed ledger (9.14.0) was backfilled by this same change rather than consumed by anyone. BACKING REQUIRES: ≥ 2 recorded instances where a reader or a downstream gate declined to treat a provisional artifact as acceptance, with the artifact and the decision both citeable. DROP: if the fields ship for one release and every consumer still reads the verdict line alone, the honest null is that the metadata is inert — the Risk-Register rank-2 outcome — and it is published as such rather than defended. |
roadmap-wall-clock-baseline |
quant | unbacked | PRE-REGISTERED 2026-08-17 (road-to-user-out-of-the-loop Phase 0 Step 3). SEPARATE FROM user-out-of-loop-baselineBY CONSTRUCTION, never merged into it: the roadmap's Goal states the two axes are deliberately not one, because a run can ask zero questions and still be slow — a single blended metric would let a contact win pay for a wall-clock loss and report the pair as progress. INSTRUMENT:src/scripts/interruption_report.tsderives elapsed time per run fromagents/runtime/.agent-chat-historytimestamps and splits it into WAITING (agent turn → the next real user turn) and WORKING (elapsed minus waiting). The join to the contact axis is the session tag: the ledger writesrun_idviaderive_session_tag, the same derivation that file writes as s. SYNTHETIC-TURN EXCLUSION is part of the definition, not a filter applied later: the harness writes task notifications and system reminders into the user role, and counting those as replies collapses every measured wait toward zero and makes the whole axis read as already-solved. The count of excluded turns is reported so the exclusion is auditable. POWER CAVEAT: identical to the sibling claim — 5 sessions measured against a 30-session request on the day of registration; window_shortis reported and a short window may not be cited as a baseline. FALSIFICATION fixed before data: (1) the QUALITY ANCHOR is the held defect rate, same as the sibling — a wall-clock win that moves it is a FAIL; (2) WORKING time, not elapsed, is the number a mechanism is judged on when the mechanism claims to remove waiting — reporting an elapsed improvement produced entirely by a faster human is the attribution error this criterion exists to block; (3) ≥ 20 recorded runs before any comparison, else UNDERPOWERED. HONEST-NULL PATH: if elapsed falls while working time does not, the recorded finding is that the change moved the human's response time and not the run's, and no autonomy claim is made from it. POST-REGISTRATION FINDING 2026-08-19 (road-to-long-horizon-execution, added AFTER registration and changing NO threshold — the ≥ 20 floor above stands exactly as written): the floor is **structurally unreachable with this instrument at default retention**, which is a different statement from "not yet reached" and has a different remedy. Timing comes only fromagents/runtime/.agent-chat-history, whose retention is DEFAULT_MAX_SESSIONS = 5 (src/scripts/chat_history.ts; chat_history.max_sessionsis unset on every settings layer, so the default is live). Five retained sessions yielded **4** timing-bearing runs on 2026-08-19, and the same file held 5 sessions at registration on 2026-08-17 — two readings two days apart, both at the cap. Waiting therefore does not fill this window; it rotates it. Backing this claim requires either a timing source that is not a rolling buffer, or an explicit retention change with its own privacy review — or this claim closes on the honest-null path above. Recorded here rather than in a roadmap because the reachability of a pre-registered floor is a property of the claim, and a reader deciding whether to wait for more data needs it at the point of the claim. The siblinguser-out-of-loop-baselineis NOT affected: its source is the committed append-only ledger, which stood at **19** of 20 on the same day and is reachable by one more recorded run.interruption_reportnow prints each axis's own N against the floor, because the single ⚠️ SHORT WINDOW banner over both axes had already produced one live misreading —runs: 21 read as the contact axis clearing the floor, when 2 of those runs carry timing and no ledger entry. |
rule-count |
quant | backed | exec:check_artefact_count_messaging -> 0 |
scope-dedup-cold-start-reduction |
quant | backed | agents/settings/contexts/cache-economy-refusals.md#Honest null — scope de-duplication is measured but |
scoped-dangle-follow-rate |
quant | unbacked | PRE-REGISTERED 2026-08-23 (road-to-skill-link-integrity-and-manifest-sync Phase 4). POPULATION, measured and reproducible: 24 dangling links from 17 surviving skills, derived by lint_handoffs –census-jsonwithis_pruned_under_scoped— the predicateinstall.tsitself applies, so the counted set cannot describe a projection the installer does not perform. METRIC: read attempts against.claude/skills/over a 30-day window, fromagents/runtime/metrics/skill-usage.jsonl. THRESHOLD, fixed now: zero attempts over a LIVE window closes this as a published null and the 24 links stay; a nonzero count promotes the fix, which is to rewrite each dangling link in the PROJECTED SKILL.md to name the slug and its pack instead of linking it — source tree untouched, using the same predicate the counter uses. INSTRUMENT STATUS: **dead, and the measurement was therefore not attempted.** Two independent reasons, both verified: (1) the store is gitignored and machine-local, so it is ABSENT in any fresh checkout, worktree, or CI run — that is the state the committed row agents/evidence/metrics/scoped-dangle-follow-rate.jsonrecords; in the maintainer's parent checkout it holds 181 records whose newest timestamp is 100 days old (2026-05-15T13:44:17.594Z), so it is stale there rather than absent. (2) **A LIVE clock would still not answer this**, which the drafting phase did not foresee: every one of those 181 records carrieskind: “exposure”, and no event in FOLLOW_KINDS (read, read_attempt, follow) is emitted anywhere in the tree — the instrument records that a skill was SHOWN, never that a link was FOLLOWED. BLOCKED ON: emitting a follow event, which is not in this roadmap. FALSIFICATION: attemptsisnulland never0wheneverinstrument_liveis false, asserted intests/scripts/scoped_dangle_window_guard.test.ts; a 0 there would be the false null this whole phase exists to prevent, and reporting one is the failure, not the finding. |
second-brain-recall-lift |
quant | backed | internal/bench/reports/second-brain-delta.json |
sequential-thinking-chain-of-thought |
qual | backed | https://arxiv.org/abs/2201.11903 (2026-08-11) |
shipped-artifacts-hidden-instruction-scanned |
qual | backed | exec:lint_agent_security -> 0 |
skill-count |
quant | backed | exec:check_artefact_count_messaging -> 0 |
skill-improvement-reflexion |
qual | backed | https://arxiv.org/abs/2303.11366 (2026-08-11) |
skill-link-census |
quant | backed | agents/evidence/metrics/skill-link-census.json#"dead_links": [] |
skill-tiering-h1-unmeasured |
quant | backed | agents/evidence/analysis/skill-tiering-matrix-arm.md#The question this arm cannot answer |
skill-tiering-h2-costs-more-by-default |
quant | backed | agents/evidence/analysis/skill-tiering-matrix-arm.md#H2 |
subagent-valid-envelope-rate |
quant | unbacked | PRE-REGISTERED 2026-08-22 (road-to-subagent-envelope-adoption Phase 1.3). BASELINE, measured before the pointer landed: **0 valid envelopes of 1,845 post-split stops**, window 2026-08-13T21:19:46Z through 2026-08-22T11:43:05Z, over agents/runtime/state/subagent-ledger/2026-08.jsonl. Verdict breakdown no_envelope1,817 ·fail28 ·ok0 ·no_message0; a further 4,543 rows carry the retiredabsentvocabulary and are excluded rather than folded in. Agent-type composition(null)1,725 ·general-purpose92 ·Explore28 — the null majority is the start-to-stop join rate of roughly 8 in 100 recorded elsewhere, and it is why no stop can be attributed to a dispatcher. METRIC:okdivided by post-split stops, reported bysrc/scripts/report_envelope_rate.ts, which prints the rate, the window bounds, the stop count and the ledger path on one line. THRESHOLD FOR THE FIRST WINDOW: greater than zero and rising. Deliberately NOT a percentage — a first window held to a high bar would fail for reasons the measurement cannot separate from the pointer's own effect, so the only thing the first window can establish is that the pointer is readable at all. POWER: the baseline denominator is 1,845; a first window below ~100 stops distinguishes nothing. FALSIFICATION: (1) a rate that stays 0 has at least three causes — the pointer is unreadable, the dominant path was misidentified, or workers on that path never emit a final assistant message — and a single rate cannot separate them, so a flat rate is reported as unresolved rather than as the pointer having failed; (2) the ledger is gitignored and machine-local, so every reading is one machine's drain traffic and no rate from it generalises; (3) a rate that rises without a contemporaneous pointer-removed arm does not establish that the pointer caused it — the historical baseline is temporally and compositionally confounded and both council seats refused it as a control. |
suggestion-capture-rate |
quant | unbacked | PRE-REGISTERED 2026-08-24 (road-to-suggestion-block-capture Phase 1.2), BEFORE any capture code exists. BASELINE: the model-carried comparator is resolved and dead — orchestration_recordcaptured 1 of 369 dispatches, and that figure "may not be cited for either direction" per its own entry, so it is the reason this instrument exists rather than a number this claim beats. THIS INSTRUMENT'S BASELINE IS ZERO BY CONSTRUCTION: the sinkagents/runtime/state/audit/suggestion-capture.jsonldoes not exist before the hook lands. METRIC: lines written to that sink divided by suggestion blocks emitted, where the denominator has a reading INDEPENDENT of the instrument under test — a contemporaneous emission log kept by the maintainer during the window. Without that independent denominator the instrument measures only itself and no rate is claimable. WINDOW: fourteen days, fixed insrc/config/suggestion-capture.json before any capture code ran, because a window whose length is chosen after the numbers are in is a window chosen to produce a number. THRESHOLD FOR THE FIRST WINDOW: greater than zero and rising. Deliberately NOT a rate figure — a first window held to a high bar fails for reasons the measurement cannot separate from the instrument's own readability, so the only thing it can establish is that capture happens at all. FALSIFICATION: a window in which the maintainer's log records blocks emitted and the sink carries zero lines DROPS this claim and parks the three consumer roadmaps' resume conditions as unsatisfiable by this instrument. SCOPE, measured rather than assumed: the payload probe covered Claude Code only (agents/evidence/analysis/suggestion-capture-probe.md), so any figure is one host on one machine and generalises to neither the other five bound platforms nor another operator's traffic. CLOCK CORRECTION 2026-08-24: no observation made before this date is admissible, and the fourteen days start at verified deployment of the FIXED instrument rather than at its commit. Until 2026-08-24 the concern declared main(now: Date = new Date())while the dispatcher callsmain(argv), so now.getTime()threw on every live turn and the instrument's own catch swallowed it — exit 0, no output, indistinguishable from a disabled hook. The consequence for THIS claim is specific rather than cosmetic: the falsification clause above drops the claim on a window whose sink carries zero lines, and applied to the broken period it would have dropped it for the wrong reason and parked three consumer roadmaps as permanently unsatisfiable on the strength of a bug.status: unbacked records that there is no evidence; it does not record which observations are admissible, which is why this correction is written here and not left to the reader. AI council 2/2. |
surgical-uninstall |
qual | backed | exec:vitest run tests/lib/json_pointers.test.ts -> 0 |
team-defect-finding-null |
quant | backed | internal/bench/reports/defect-finding.json#honest_null |
thin-inject-delivery-equivalence |
quant | backed | internal/bench/reports/thin-inject-2026-08-23.md#endpoints |
unattended-demotion-gate |
quant | resolved-null | PRE-REGISTERED 2026-08-19 (road-to-long-horizon-execution Phase 4.2, sequencing UOTL Phase 7.3). REGISTERED BEFORE THE CAPABILITY EXISTS, which is the point and is stated rather than implied: at registration time NO unattended run has occurred and none can, because the headless spawn is deliberately unbuilt (unattended_guard.ts§ "Why the spawn is not in this file") and the budget defaults to both ceilings zero, which disables the lane rather than permitting it. So this entry cannot have been written around a number already in hand. THRESHOLD, fixed now: the rework rate of PRs produced by unattended runs, measured over the 14 days after each merges, must not exceed the same-window rate for attended PRs; a breach flipsmax_usd/max_tokensback to 0 in the same change that reports the number. REWORK is defined before any data exists, because a metric defined after the fact is chosen: a follow-up commit touching a file the run's PR touched, within 14 days of merge, excluding (a) commits by the same run continuing planned roadmap work, (b) pure dependency bumps, (c) reverts of an unrelated change that merely collide. POWER: at least 10 unattended PRs and 10 attended PRs in the comparison window, else UNDERPOWERED and no claim either way — a two-PR sample producing a favourable ratio is the failure this line exists against. FALSIFICATION: (1) an unattended PR that a human had to substantially rewrite counts as rework even when no commit touched the same file, and that case is recorded by hand rather than dropped because the mechanical definition missed it; (2) the comparison is rate-vs-rate, never absolute counts, since the two populations will not be the same size; (3) a rate that is LOWER for unattended runs is reported as-is and is NOT used to argue for widening the lane — this gate can close a lane, never open one. HONEST-NULL PATH: if the lane never runs (the spawn stays unbuilt, or the budget stays at zero), the recorded finding is that the gate was pre-registered and never had data, and the capability is closed rather than left indefinitely pending — the same D-5 shape this roadmap opens by naming. **HONEST-NULL PATH TAKEN 2026-08-19 — the SAME DAY it was registered, and that is stated rather than rounded** (R2 round 1, finding 2 caught this entry claiming "one day after"; registration above is dated 2026-08-19 too, so the elapsed time is zero). A pre-registration closed on its own registration date deserves the suspicion it attracts, so here is why it is not a threshold written around a result: the entry was registered when the spawn was DEFERRED, and closed when the spawn became REFUSED. Nothing measured moved in between — a decision did, and it is recorded with its council and its reasoning. The threshold never met data in either state. The lane will not run: the headless spawn is no longer "deliberately unbuilt" but a published refusal (road-to-long-horizon-execution 4.0, AI council 2026-08-19), and the two acceptance criteria that depended on it are cancelled WILL-NOT-MEASURE. So the recorded finding is exactly the one this path fixed in advance: **the gate was pre-registered, never had data, and the capability is closed** — zero unattended PRs against the ≥ 10-vs-10 power floor, which is not a null result but an absent population, and the entry says which. Nothing here is read as evidence in either direction, and specifically not as "unattended runs are safe": an unrun lane has no rework rate. The threshold, the rework definition and the power floor are left EXACTLY as registered so that a future roadmap reopening the capability inherits a bar written before anyone knew the answer; reopening requires that new roadmap, never a re-reading of this entry. The reopen trigger is 4.0's: the first checkpoint written by a real dying run. STATUSresolved-null, not unbacked (R2 round 2, finding 3): the ledger already has the terminal status for exactly this state, and leaving a closed question inside the documented-debt inventory would inflate the unbacked count and leave it looking pending — the indefinite-pending shape this entry argues against, reproduced by the entry that argues against it. |
user-out-of-loop-baseline |
quant | unbacked | PRE-REGISTERED 2026-08-17 (road-to-user-out-of-the-loop Phase 0 Step 3 — registered BEFORE the instrument had recorded a single observation, and deliberately BEFORE any Phase 1 mechanism ships, so the baseline cannot be read after the change it is meant to judge). INSTRUMENT: the interruption-ledger concern (src/scripts/hooks/interruption_ledger_hook.ts, stopslot, capture-only) writes one{run_id, turn, kind, class, roadmap}line per turn toagents/runtime/state/interruptions.jsonl; src/scripts/interruption_report.ts reads it. DEFINITION fixed before data, and it is three classes rather than two on purpose: a CONTACT is a turn whose closing paragraph either ends in a question (ask) or yields the decision without one (handback). Counting only ?would score this package's own preferred hand-back shape as zero contacts and make the metric flatter the design — the Risk-6 failure the roadmap names. POWER CAVEAT recorded at registration, not discovered later: the rolling chat history is a buffer, not an archive — measured the day this was written it held **5 sessions, all from one day**, against the 30-session conformance window the step asks for. The report therefore reportssessions_foundnext towindow_requestedand flagswindow_short; a number computed over a short window may not be cited as an N-session baseline in either direction. FALSIFICATION fixed before data: (1) the QUALITY ANCHOR is the held defect rate — a contact reduction that moves the defect rate is a FAIL regardless of its own number, and the two are never reported apart; (2) the baseline must rest on ≥ 20 recorded runs before any post-change comparison is made, and a comparison against fewer is reported UNDERPOWERED, never as a win; (3) a run present in the ledger but not the history, or the reverse, is reported with the missing axis null and is never scored as zero — scoring an unmeasured run as zero contacts is the arithmetic that would manufacture the result. HONEST-NULL PATH: if the median does not move, or moves only with the defect rate, the null is published and Phase 1's mechanisms are judged on it rather than the claim being re-scoped to a family that happened to win. |
utilization-window-decidability |
comparative | unbacked | PRE-REGISTERED 2026-07-12 (road-to-feedback-8.11-2 Phase 0 — no goalpost-moving after the numbers land; criteria at docs/design/utilization-window-criteria.md). Floor fixed BEFORE data: >=100 task boundaries AND >=2 hosts (or the documented degraded form) AND >=45 elapsed days; decision rules D1 (loaded-never-consulted -> retirement-candidate list), D2 (consulted-never-applied <10% applied-ratio at >=5 consultations -> trigger-review queue), D3 (above floor -> >=1 named decision per kind or a recorded why-not), D4 (below floor after one extension -> honest null, lifecycle/ledger gates stay closed). Kernel + safety floors exempt by construction. |
wedge-hollow-detection |
quant | backed | internal/bench/orchestration/pv-a3-results.md#token_delta_provenance: measured |
worker-capsule-trigger-arm |
comparative | unbacked | PRE-REGISTERED 2026-08-09 (road-to-worker-generation-recycling Phase 1.4 — registered BEFORE the first shadow capsule is read; the mechanism ships shadow-only, so no capsule has been scored at registration time). CAPSULE-QUALITY RUBRIC, fixed here, five binary criteria scored 0-5 per capsule — (1) remaining[]names every open item the task still needs, no silent drops; (2)decisions[]names each choice a successor would otherwise silently re-open; (3)assumptions[]is non-empty and every entry carries a resolvingbasisref; (4) everydone[] ref resolves to a real file/line; (5) a successor briefed on the ORIGINAL brief plus the capsule alone takes a first action that neither repeats completed work nor asks for a re-brief. ADOPTION MARGIN, fixed BEFORE data: an arm is adopted only if, on paired samples from the same runs, it fires at a median of >= 2 steps earlier AND its capsules score >= 4/5 on the rubric with no regression against the other arm; an arm that wins on earliness while dropping below 4/5 is NOT adopted, because an earlier bad capsule is worse than a later good one. Sample floor: >= 30 shadow capsules with BOTH trigger points recorded (watermark_step, saturation_step, trigger_arm_earlieron theorchestration_recordline). Instrument:src/scripts/_lib/capsule_trigger.ts (compareTriggers, earlierArm), term-frequency only, no embeddings. HONEST-NULL consequence, pre-authorised: BOTH arms losing (neither reaches 4/5, or the margin is not met) is a publishable result that closes the mechanism as default-off — it is the expected-value outcome given the standing orchestration-observed-dispatch-cost null, and it must be cheap to record. Token delta is reported as a pair with quality and is explicitly NOT the claim. |
5. Verify it yourself
Section titled “5. Verify it yourself”On a fresh checkout, reproduce the claims above:
task check-claims # every markered public claim binds to resolvable evidencetask check-refs # no broken internal referencestask check-skill-gaps # every logged known-limit cites a real witness testtask check-comparison # every comparison-table "our evidence" pointer resolves./scripts-run src/scripts/skill_eval_coverage # behavioural-eval coverage, per tier./scripts-run src/scripts/skill_eval_coverage --check # the ratchet: coverage may not droptask build-proof-check # this page is in sync with its sourcesIf a claim ever loses its binding, or this page drifts from the ledger, CI goes red. Reproducibility is the proof.