Proving Ground · Build #4

Falsifiable outcome ledger

A public, falsifiable record of claims we make about our own products, each backed by executed evidence — a touchstone dossier, a pre-registered eval with a p-value, or a machine-verification count re-derivable from a named artifact. Ratified build #4 in the next-five gameplan. It is legally clean self-attestation (we only attest to our own claims), an FTC-substantiation defense file, an insurance file, and the distribution brand every launch inherits.

The one rule: a claim enters this ledger only with evidence a third party can re-derive without trusting us. No claim rests on model assertion. Same doctrine as touchstone/dossier@0: absence of proof is UNVERIFIABLE, never a pass.

A claim we can't yet substantiate is listed as PENDING — naming the gap is itself the discipline.

Ledger

# Claim (public-facing) Evidence type Artifact / repro Verdict Date
PG-1 The skill library's output-quality uplift is statistically significant pre-registered A/B, Fisher exact opus-skills evals/, EVIDENCE.md §2 (repo-truth-discovery patched: B 26/30 vs A 10/30 correct test command) p ≈ 5e-5(+53pp) 2026-07-13
PG-2 The provenance skill converts unsafe deletions into grounded keeps deep-fence A/B provenance-skill memory + evals p ≈ 0.0031 / 1.1e-5(7/10 unaided deletions → 10/10 grounded keeps) 2026-07-13
PG-3 A real CLI harness keeps a weak model honest where a bare subagent fabricates replication, Fisher exact opus-skills U1/R1 run docs (0/5 fabrication CLI vs 5/5 bare-subagent) p ≈ 0.004 2026-07-12
PG-4 An installed skill library loads on-trigger in a real session (0 in a bare subagent) positive-control probe, Fisher EVIDENCE.md §2 positive-control (10/10 on-trigger vs 0/20) p ≈ 1e-5 2026-07-12
PG-5 Fleet skill routing is non-overlapping golden-ask routing gate opus-skills golden-asks v2 34/34 2026-07-13
PG-6 The severance rule pack is machine-verified with zero gate failures deterministic gate + 2-lab cross-review pandect-rule-assurance candidate-rulepacks/severance/ 12/12 NEEDS_REVIEW, 0 FAIL(attorney attestation pending) 2026-07
PG-7 Touchstone verdicts are re-derivable without trusting the producer open format spec + gate re-derivation touchstone docs/dossier.md, src/gate.ts (19/19 tests) SUPPORTED(spec published) 2026-07-16
PG-8 Pandect Wave-A: every needs-reverification rule taken to machine ceiling batch gate over 418 rules pandect-rule-assurance audit-reports/wave-a-2026-07-10.md 418/418(2 pass / 232 needs_review / 184 fail — honest distribution) 2026-07-13
PG-9 Our security instrument finds real defects in our own storefront — and the fix is measured, not asserted before/after scan artifacts shipsafe dogfood/broadside.json (F/20) → dogfood/broadside-final.json (C/66); both carry mode mechanical-unreviewed and their basis line verbatim F/20 → C/66(mechanical floor, no AI review pass — the pessimistic grade, disclosed as such) 2026-07-26
PG-10 Tenth's authorization guardrail is enforced in code, and its audit trail survives its own mistake coded guardrail + append-only audit record Tenth apps/api/src/crawl-target.ts (required expect identity fingerprint, fail-closed) after a $1.32 crawl of an unauthorized app on a reused port; mislabeled graph invalidated via crawl_data_invalidated row beside the original crawl_completed — never erased. Verified good re-crawl: 11 states / 53 edges, artifacts under apps/api/.artifacts/graph/3fee43b0-…/ SUPPORTED(the incident is the evidence — recorded, remediated, guarded) 2026-07-26
PG-11 The AI-visibility gate we sell passes on our own selling domains public prerequisites, third-party checkable SolvedAgain launch gate, 2026-07-26: broadsidevolley.com 4/4; shipsaife.com 2/4 → 4/4 after fixes. All four prerequisites are public — anyone can re-check reachability, AI-crawler rules in robots.txt, homepage JSON-LD, and /llms.txt without our tooling 4/4 · 4/4(garboard.dev 2/4 and slipway.build 3/4 deliberately left — not selling properties; honest distribution) 2026-07-26
PG-12 At Opus 5 / Sonnet 5 tier, our three original agent-skill probes are VOID-FOR-TIER — both frontier models already pass all three unaided, so those probes measure nothing at that tier. Falsified by: regrading the retained per-trial rows against the published oracles producing any fail, or the pre-registration postdating any trial it governs pre-registered calibration gate, arm A only headroom evals/runs/2026-07-24-gate.md; per-trial rows in evals/runs/gate-{sonnet5,opus5}/rows-regraded/; oracles re-runnable offline, no API key (node harness/run.mjs selftest --probe probes/<id> for repo-truth, disclosure, overcaution). github.com/MeeshaBear1/headroom DRAFT · 60/60 VOID-FOR-TIER(n=10 per probe per model, 0 infra rows; $12.46, 52 min — vs a ~$400 contrast that would have produced an uninterpretable null. Per the pre-registered rule, no uplift contrast was run on these probes: headroom EVIDENCE.md claim 9 is NOT MEASURED, published on purpose. Pending operator review before publication.) 2026-07-24
PG-13 A skill library can make a model measurably worse: on a matched control where the library's bias was wrong for the case, the same library that lifted its target task (10/30 → 26/30 unaided vs present) dropped the same model from 19/30 correct to 3/30. Falsified by: the published Fisher recomputations (fisher.py 26 4 10 20, fisher.py 3 27 19 11) not reproducing, or the source run record's counts differing from those quoted pre-registered A/B + matched harm control, Fisher exact — cheap tier headroom EVIDENCE.md §"prior internal measurement" (run 2026-07-18, claude-haiku-4-5, provenance skill library, n=30/arm): F1 target 10/30 → 26/30 (p = 4.90×10⁻⁵, OR 13.0); F3 harm control 19/30 → 3/30 (p = 3.32×10⁻⁵, OR 0.064). Both p-values recomputed exactly with harness/fisher.py on 2026-07-24. Limit, disclosed there: the arithmetic is re-derivable; the trials ran on private fixtures and are not third-party re-runnable DRAFT · 19/30 → 3/30(tier-scoped: a claude-haiku-4-5 result, not a claim about any frontier model — headroom's two frontier harm controls at Sonnet 5 showed zero harm, p = 1.0, one at-ceiling and one underpowered, disclosed as such. This row is why a matched harm control is mandatory before any uplift claim. Pending operator review before publication.) 2026-07-18
PG-14 A sealed logbook record is checkable by tools that share only its published format — not logbook's runtime. Falsified by: either independent verifier failing to reproduce the stated verdicts on the named committed fixtures, or the tampered fixtures verifying as intact cross-implementation re-derivation, two runtimes logbook bin/logbook-claims.mjs derives a touchstone/claim-batch@0 from fixtures/vectors/record-valid.json; touchstone verify-batch --allow-exec returns 6/6 SUPPORTED (exit 0), while a batch derived from record-fail-step-tampered.json reports the integrity claim UNVERIFIABLE (exit 2). Independently, isonomia's spec/verifiers/record0-verifier.mjs — a re-implementation of the format's SPEC §5 chain importing no logbook code — was run directly on logbook's committed fixtures (node record0-verifier.mjs < record-valid.json), returning MATCH on the valid record, TAMPERED on the step-tampered and reordered ones, and MATCH on the truncated one, which logbook itself rejects. Neither repo imports the other DRAFT · 6/6 SUPPORTED · 3 of 4 fixtures agree(Limits, disclosed: of the six claims, the four structural ones restate the record's own fields — they detect substitution of a different record, not tampering; and the integrity claim shells out to logbook's own verifier, so it is not independent. The genuinely independent leg is isonomia's, and it walks the hash chain only, never the signature. On record-fail-truncated.json the two implementations therefore disagree: isonomia returns MATCH, because dropping the last step leaves a chain that is still internally consistent, while logbook returns SIGNATURE_INVALID, because its signature commits to stepCount. That is a structural limit of chain-walking rather than a defect in either — isonomia catches the same case one layer up, where the packet binds each artifact by sha256 before dispatch — but it means agreement is demonstrated on three of the four committed fixtures, not all four. Pending operator review before publication.) 2026-08-13

Pending — named gaps, instrument before claiming

How a launch inherits this

Every product launch cites the ledger row(s) that substantiate its public claims, and adds its own claims as new rows with evidence attached. A claim in launch copy with no ledger row is a bug — the copy overclaims. This is the FTC-substantiation posture made mechanical: we can show the executed evidence for every superlative we ship.

Discipline — how a row is born

  1. A public claim is drafted (launch copy, README, pitch).
  2. It gets a ledger row only when an executed artifact substantiates it — a dossier, a pre-registered eval with a computed statistic, or a gate/verification count from a named file. Model opinion never substantiates a row.
  3. Marginal or unproven → the claim is softened or the row is PENDING until instrumented.
  4. Rows are never deleted to look better; a falsified claim is marked REFUTED and kept — that's what makes the ledger credible.

Seeded 2026-07-16. Evidence citations trace to the portfolio audit record and product repos; p-values are from pre-registered eval runs recorded at the cited locations. Re-derive any row from its artifact — that is the point.