Proving Ground · Build #4
A public, falsifiable record of claims we make about our own products, each backed by executed evidence — a touchstone dossier, a pre-registered eval with a p-value, or a machine-verification count re-derivable from a named artifact. Ratified build #4 in the next-five gameplan. It is legally clean self-attestation (we only attest to our own claims), an FTC-substantiation defense file, an insurance file, and the distribution brand every launch inherits.
The one rule: a claim enters this ledger only with evidence a third
party can re-derive without trusting us. No claim rests on model assertion. Same doctrine
as touchstone/dossier@0: absence of proof is UNVERIFIABLE, never
a pass.
A claim we can't yet substantiate is listed as PENDING — naming the gap is itself the discipline.
| # | Claim (public-facing) | Evidence type | Artifact / repro | Verdict | Date |
|---|---|---|---|---|---|
| PG-1 | The skill library's output-quality uplift is statistically significant | pre-registered A/B, Fisher exact | opus-skills evals/, EVIDENCE.md §2 (repo-truth-discovery patched: B 26/30 vs A 10/30 correct test command) |
p ≈ 5e-5(+53pp) | 2026-07-13 |
| PG-2 | The provenance skill converts unsafe deletions into grounded keeps | deep-fence A/B | provenance-skill memory + evals | p ≈ 0.0031 / 1.1e-5(7/10 unaided deletions → 10/10 grounded keeps) | 2026-07-13 |
| PG-3 | A real CLI harness keeps a weak model honest where a bare subagent fabricates | replication, Fisher exact | opus-skills U1/R1 run docs (0/5 fabrication CLI vs 5/5 bare-subagent) | p ≈ 0.004 | 2026-07-12 |
| PG-4 | An installed skill library loads on-trigger in a real session (0 in a bare subagent) | positive-control probe, Fisher | EVIDENCE.md §2 positive-control (10/10 on-trigger vs 0/20) | p ≈ 1e-5 | 2026-07-12 |
| PG-5 | Fleet skill routing is non-overlapping | golden-ask routing gate | opus-skills golden-asks v2 | 34/34 | 2026-07-13 |
| PG-6 | The severance rule pack is machine-verified with zero gate failures | deterministic gate + 2-lab cross-review | pandect-rule-assurance candidate-rulepacks/severance/ |
12/12 NEEDS_REVIEW, 0 FAIL(attorney attestation pending) | 2026-07 |
| PG-7 | Touchstone verdicts are re-derivable without trusting the producer | open format spec + gate re-derivation | touchstone docs/dossier.md, src/gate.ts (19/19 tests) |
SUPPORTED(spec published) | 2026-07-16 |
| PG-8 | Pandect Wave-A: every needs-reverification rule taken to machine ceiling | batch gate over 418 rules | pandect-rule-assurance audit-reports/wave-a-2026-07-10.md |
418/418(2 pass / 232 needs_review / 184 fail — honest distribution) | 2026-07-13 |
| PG-9 | Our security instrument finds real defects in our own storefront — and the fix is measured, not asserted | before/after scan artifacts | shipsafe dogfood/broadside.json (F/20) → dogfood/broadside-final.json (C/66); both carry mode mechanical-unreviewed and their basis line verbatim |
F/20 → C/66(mechanical floor, no AI review pass — the pessimistic grade, disclosed as such) | 2026-07-26 |
| PG-10 | Tenth's authorization guardrail is enforced in code, and its audit trail survives its own mistake | coded guardrail + append-only audit record | Tenth apps/api/src/crawl-target.ts (required expect identity fingerprint, fail-closed) after a $1.32 crawl of an unauthorized app on a reused port; mislabeled graph invalidated via crawl_data_invalidated row beside the original crawl_completed — never erased. Verified good re-crawl: 11 states / 53 edges, artifacts under apps/api/.artifacts/graph/3fee43b0-…/ |
SUPPORTED(the incident is the evidence — recorded, remediated, guarded) | 2026-07-26 |
| PG-11 | The AI-visibility gate we sell passes on our own selling domains | public prerequisites, third-party checkable | SolvedAgain launch gate, 2026-07-26: broadsidevolley.com 4/4; shipsaife.com 2/4 → 4/4 after fixes. All four prerequisites are public — anyone can re-check reachability, AI-crawler rules in robots.txt, homepage JSON-LD, and /llms.txt without our tooling |
4/4 · 4/4(garboard.dev 2/4 and slipway.build 3/4 deliberately left — not selling properties; honest distribution) | 2026-07-26 |
| PG-12 | At Opus 5 / Sonnet 5 tier, our three original agent-skill probes are VOID-FOR-TIER — both frontier models already pass all three unaided, so those probes measure nothing at that tier. Falsified by: regrading the retained per-trial rows against the published oracles producing any fail, or the pre-registration postdating any trial it governs | pre-registered calibration gate, arm A only | headroom evals/runs/2026-07-24-gate.md; per-trial rows in evals/runs/gate-{sonnet5,opus5}/rows-regraded/; oracles re-runnable offline, no API key (node harness/run.mjs selftest --probe probes/<id> for repo-truth, disclosure, overcaution). github.com/MeeshaBear1/headroom |
DRAFT · 60/60 VOID-FOR-TIER(n=10 per probe per model, 0 infra rows; $12.46, 52 min — vs a ~$400 contrast that would have produced an uninterpretable null. Per the pre-registered rule, no uplift contrast was run on these probes: headroom EVIDENCE.md claim 9 is NOT MEASURED, published on purpose. Pending operator review before publication.) | 2026-07-24 |
| PG-13 | A skill library can make a model measurably worse: on a matched control where the library's bias was wrong for the case, the same library that lifted its target task (10/30 → 26/30 unaided vs present) dropped the same model from 19/30 correct to 3/30. Falsified by: the published Fisher recomputations (fisher.py 26 4 10 20, fisher.py 3 27 19 11) not reproducing, or the source run record's counts differing from those quoted |
pre-registered A/B + matched harm control, Fisher exact — cheap tier | headroom EVIDENCE.md §"prior internal measurement" (run 2026-07-18, claude-haiku-4-5, provenance skill library, n=30/arm): F1 target 10/30 → 26/30 (p = 4.90×10⁻⁵, OR 13.0); F3 harm control 19/30 → 3/30 (p = 3.32×10⁻⁵, OR 0.064). Both p-values recomputed exactly with harness/fisher.py on 2026-07-24. Limit, disclosed there: the arithmetic is re-derivable; the trials ran on private fixtures and are not third-party re-runnable |
DRAFT · 19/30 → 3/30(tier-scoped: a claude-haiku-4-5 result, not a claim about any frontier model — headroom's two frontier harm controls at Sonnet 5 showed zero harm, p = 1.0, one at-ceiling and one underpowered, disclosed as such. This row is why a matched harm control is mandatory before any uplift claim. Pending operator review before publication.) | 2026-07-18 |
| PG-14 | A sealed logbook record is checkable by tools that share only its published format — not logbook's runtime. Falsified by: either independent verifier failing to reproduce the stated verdicts on the named committed fixtures, or the tampered fixtures verifying as intact | cross-implementation re-derivation, two runtimes | logbook bin/logbook-claims.mjs derives a touchstone/claim-batch@0 from fixtures/vectors/record-valid.json; touchstone verify-batch --allow-exec returns 6/6 SUPPORTED (exit 0), while a batch derived from record-fail-step-tampered.json reports the integrity claim UNVERIFIABLE (exit 2). Independently, isonomia's spec/verifiers/record0-verifier.mjs — a re-implementation of the format's SPEC §5 chain importing no logbook code — was run directly on logbook's committed fixtures (node record0-verifier.mjs < record-valid.json), returning MATCH on the valid record, TAMPERED on the step-tampered and reordered ones, and MATCH on the truncated one, which logbook itself rejects. Neither repo imports the other |
DRAFT · 6/6 SUPPORTED · 3 of 4 fixtures agree(Limits, disclosed: of the six claims, the four structural ones restate the record's own fields — they detect substitution of a different record, not tampering; and the integrity claim shells out to logbook's own verifier, so it is not independent. The genuinely independent leg is isonomia's, and it walks the hash chain only, never the signature. On record-fail-truncated.json the two implementations therefore disagree: isonomia returns MATCH, because dropping the last step leaves a chain that is still internally consistent, while logbook returns SIGNATURE_INVALID, because its signature commits to stepCount. That is a structural limit of chain-walking rather than a defect in either — isonomia catches the same case one layer up, where the packet binds each artifact by sha256 before dispatch — but it means agreement is demonstrated on three of the four committed fixtures, not all four. Pending operator review before publication.) |
2026-08-13 |
audit_batch.py report emits the batch report.Every product launch cites the ledger row(s) that substantiate its public claims, and adds its own claims as new rows with evidence attached. A claim in launch copy with no ledger row is a bug — the copy overclaims. This is the FTC-substantiation posture made mechanical: we can show the executed evidence for every superlative we ship.
Seeded 2026-07-16. Evidence citations trace to the portfolio audit record and product repos; p-values are from pre-registered eval runs recorded at the cited locations. Re-derive any row from its artifact — that is the point.