The four headline claims, each with pinned source evidence, direct downloads, reproduce commands, and honest scope boundaries.
Single source of truth for the four headline claims. All four were independently fact-checked on 2026-07-04; the TEP ownership and current-profile scope were rechecked on 2026-08-01, and its merged v8.1 paper and replay evidence were rechecked on 2026-08-03. On 2026-08-23 the tep-far0 claim ALONE was substantially rewritten around a new held-out Rieth-2017 result and its figures re-measured from the committed raw JSON; the other three claims were NOT rechecked on that date, which is why fact_checked still reads 2026-08-03. On 2026-08-24 one tep-far0 gap was CLOSED rather than reworded: component B's detection latency on the masked trio was measured and is no longer absent, which moved both evidence branch heads and forced a re-pin of that claim's SHAs, its paper page count and its paper digests. Still only tep-far0; fact_checked is deliberately unchanged. Corrections are recorded in CLAIMS.md. Evidence is pinned to immutable commit SHAs. Source links marked "private" require cubie-tf / cubie-research repository access (both proprietary); public downloads are served from lib.trustfortress.ai. Numbers quantifying repo state are provenance snapshots — re-run the reproduce commands to re-measure. Library evidence under email-gated IP prefixes is labelled "email-gated" on this public surface and requires authorized Cloudflare Access before bytes are served.
TEP detector evidence: held-out Rieth-2017 result — 0 of 500 fault-free runs alarmed, all 20 faults detected — plus superseded in-sample v8.1 replay and LLM comparison Scoped
Scope: A held-out result, plus retained historical in-sample material that it does NOT vindicate. The headline 0/500 and 20/20 figures come from the Rieth-2017 TESTING splits, scored with coefficients, item selection and thresholds derived from TRAINING runs only and frozen, hashed and committed before any testing byte was read; the fit tool refuses any input whose basename contains 'Testing', and training is split three ways so no quantity is estimated and evaluated on the same runs. This is NOT a pristine first-exposure estimate: the fault-free testing file has now been scored three times across protocol versions and the faulty file twice, and the decision to restrict the residue channel to three items was PROMPTED BY (though not justified by) an observed false alarm in a wider eighteen-item variant. The strongest figure in this record that is test-blind in every fitted quantity remains component A's own — 0 false alarms with 17 of 20 faults at run-level 1.0000. Component A is NOT fully test-blind in its structure, and this is stated rather than implied: its seam layout is not canonical but the output of a search whose acceptance test scored candidates on the Braatz TEST traces (layoutSha256 fd1a3acd… resolves exactly to data/tep/layouts/_perfect_3wins_seed73_cusum.json), and the layout is load-bearing because it decides which variable pairs form the twelve conditional seams. The bias this could introduce is measured, not assumed: five profiles were fit on FIVE DISTINCT layouts and all five were scored on the held-out splits, and under every one of them the union is 0/500 fault-free and 20/20 at run level, with component A missing exactly IDV-3/9/15 every time. The property that search optimised for on Braatz — masked-trio detection at 1.0 — transfers to Rieth at 0.0 under all five. The detector reported here is a UNION OF TWO COMPONENTS and is a different detector from the legacy CUSUM/layout configuration; the 100/100/100 @ d00 FAR=0.000% result remains an in-sample calibration identity and is retained below as superseded historical provenance, not as support for this claim. Component B is a Python reference implementation, not the shipped Rust detector.
Evidence, downloads & reproduce steps →