TEP detector evidence: held-out Rieth-2017 result — 0 of 500 fault-free runs alarmed, all 20 faults detected — plus superseded in-sample v8.1 replay and LLM comparison
Scoped
Scope boundary: A held-out result, plus retained historical in-sample material that it does NOT vindicate. The headline 0/500 and 20/20 figures come from the Rieth-2017 TESTING splits, scored with coefficients, item selection and thresholds derived from TRAINING runs only and frozen, hashed and committed before any testing byte was read; the fit tool refuses any input whose basename contains 'Testing', and training is split three ways so no quantity is estimated and evaluated on the same runs. This is NOT a pristine first-exposure estimate: the fault-free testing file has now been scored three times across protocol versions and the faulty file twice, and the decision to restrict the residue channel to three items was PROMPTED BY (though not justified by) an observed false alarm in a wider eighteen-item variant. The strongest figure in this record that is test-blind in every fitted quantity remains component A's own — 0 false alarms with 17 of 20 faults at run-level 1.0000. Component A is NOT fully test-blind in its structure, and this is stated rather than implied: its seam layout is not canonical but the output of a search whose acceptance test scored candidates on the Braatz TEST traces (layoutSha256 fd1a3acd… resolves exactly to data/tep/layouts/_perfect_3wins_seed73_cusum.json), and the layout is load-bearing because it decides which variable pairs form the twelve conditional seams. The bias this could introduce is measured, not assumed: five profiles were fit on FIVE DISTINCT layouts and all five were scored on the held-out splits, and under every one of them the union is 0/500 fault-free and 20/20 at run level, with component A missing exactly IDV-3/9/15 every time. The property that search optimised for on Braatz — masked-trio detection at 1.0 — transfers to Rieth at 0.0 under all five. The detector reported here is a UNION OF TWO COMPONENTS and is a different detector from the legacy CUSUM/layout configuration; the 100/100/100 @ d00 FAR=0.000% result remains an in-sample calibration identity and is retained below as superseded historical provenance, not as support for this claim. Component B is a Python reference implementation, not the shipped Rust detector.
Claim
HELD-OUT RESULT (2026-08-23, protocol v3). On the held-out Rieth-2017 testing splits (Harvard Dataverse doi:10.7910/DVN/6C3JR1), a union of two frozen detectors alarms on 0 of 500 fault-free runs and reaches run-level detection 1.0000 on all twenty faults, 500 runs each, zero right-censored. For 0 alarmed runs out of 500 the exact one-sided 95% binomial upper bound is 1 - 0.05^(1/500) = 0.005974 (0.5974% per run); the supported wording is observed 0/500, NOT FAR = 0. Component A is the pre-existing marginal-residual detector, which carries 17 faults at 1.0000 with shatters_total=0 and far_proxy=0.0 across 480,000 fault-free rows. Component B is a new three-item full-conditional residue channel (artifact SHA-256 0e88662d67b8d459552be5bf160d9cfa40c7648635fd90afc436e9a99eed0f13) asked only for IDV-3/9/15, the three disturbances the decentralised controller masks; each of 52 measurements is predicted from all other measurements at the same instant plus 16 lags of all 52 (885 regressors), divided by fitted sigma, and reduced to four window moments over a 360-sample window at stride 8. Measured alone on the faulty testing file, component B reaches 1.0000 on 19 of 20 faults and 0.0020 on IDV-4, which component A carries at 1.0000; its worst-run post-onset margins on the masked trio are x1.982 (IDV-3), x4.166 (IDV-9) and x5.537 (IDV-15) against a worst clean window of x0.4376 across all 500 fault-free runs. Every fault has at least one component at 1.0000, so the union requires no run-level alignment between components. THIS SUPERSEDES THE PRIOR FAULT-9 FAILURE: the corrected 50-active-cell profile's grouped-Training fault 9 = 4/500 (0.8%) missed its predeclared 1% floor, whereas fault 9 is now detected on 500 of 500 held-out runs. It does NOT vindicate the legacy 100/100/100 figure, which belongs to a different detector and remains an in-sample CUSUM calibration identity (h = 1.5 x observed d00 peak, so zero false alarms on the calibration set is true by construction). SUPERSEDED HISTORICAL MATERIAL FOLLOWS. The merged v8.1 research paper at cubie-research commit bd51e5e reports the complete recovered seed-19 replay over d00/d03/d09/d15: 2,400/2,400 fault-active alarms, 125/1,440 clean-labelled alarms (all 125 before the declared fault onset), d00 0/960 alarms, and 3,715/3,840 correct rows (96.745%). On the common 192-window, 52-channel schedule, Cubie scored 184/192 (95.833%); strict Opus 4.7 scored 96/192 (50.000%), strict Opus 4.8 85/192 (44.271%), and strict Fable 5 105/192 (54.688%). With provider-returned adaptive-thinking summaries, Opus 4.7 scored 102/192 (53.125%), Opus 4.8 86/192 (44.792%), and Fable 5 102/192 (53.125%); Cubie is unchanged because it is local and deterministic. The May-era search also reproduced 60/60 saturated searches and the historical 100%/100%/100% with d00 0/960 control, but that layout and threshold policy were selected on the exposed traces, so it is not an out-of-sample false-alarm or sensitivity estimate. The corrected 50-active-cell schema-v3 reference has five-fold whole-run grouped-Training estimates of fault 3 = 463/500 (92.6%), fault 15 = 148/500 (29.6%), and fault 9 = 4/500 (0.8%); fault 9 failed the predeclared 1% floor. No complete sealed provider, exact float/compiler/product replay, untouched confirmation set, hardware benchmark, soak, or production authority exists.
Direct downloads (public — lib.trustfortress.ai)
Merged v8.1 TEP reproduction and evidence pack pinned to cubie-research bd51e5e. Contains the credential-safe Anthropic harness, strict and adaptive Opus 4.7/4.8/Fable 5 receipts, complete Cubie replay receipts, certification, and paper source/PDF.
⬇ Downloadsha256 613179f7c24afda40ea2237f4b9d9a72a07b65c71b069ae16761a2c641eb005b V8.1 TEP whitepaper PDF - merged research account of the recovered full-row replay, common 52-channel/960-row LLM comparison, token/time/cost receipts, and corrected Training-only profile boundary.
⬇ Downloadsha256 80c017a2f9d8d6142fd964d6e47964d9d0973e1634d181014c78fb22bc8a7e75 Historical pre-retest May recovery subset pinned to cubie-research 3a9950f. Retained for provenance; superseded for current replay and paper results by the bd51e5e v8.1 pack above.
⬇ Downloadsha256 43ca2f059655b3d334aaae24a87e1a8aa25b161bfba556b31d07f5d359a7bbd1 Original-byte May 28 strict Opus 4.7/4.8 result JSON.
⬇ Downloadsha256 ae37ee3ae03b687f41c10af6d5bfcb72281c4180352161bc52c3c4c20422b4e6 Original-byte May 29 Cubie / Opus 4.8 head-to-head summary JSON.
⬇ Downloadsha256 8665dae889583ad4b8b4ee36ab3be96d7ea16ed5ef087dd3395e0f483eafd9be Historical V6 TEP whitepaper PDF - retained for provenance and superseded by the v8.1 research paper for current replay and readiness claims. States the May same-trace 100%/100%/100% and d00 FAR = 0.000% calibration control.
⬇ Downloadsha256 8569EC5FF26A77E3B55591E11FF8A75DB26EBEB82B523E94F376696AC60B9B05
Source evidence (pinned)
PROTOCOL V3 SETTLED RESULT — the held-out record: predeclaration, raw unedited result JSON for both testing splits, and the union accounting. Read PREDECLARED_PROTOCOL.md section 7 for the exposure disclosure.
🔒 Sourcecubie-tf/docs/audit/tep_retest_2026-08/protocol_v3/RESULTS.md@d3b86dcdprivate repo Predeclared protocol v3 — fixes the artifact SHA-256, alarm rule, numeric floors, outcome classes and exact scoring commands BEFORE the scorer was pointed at a testing file. Section 7 carries the exposure accounting.
🔒 Sourcecubie-tf/docs/audit/tep_retest_2026-08/protocol_v3/PREDECLARED_PROTOCOL.md@d3b86dcdprivate repo tools/tep_conditional_channel.py — the full-conditional residue channel: fit and score subcommands. fit REFUSES any input whose basename contains 'Testing', which is what makes the train/test separation structural rather than procedural.
🔒 Sourcecubie-tf/tools/tep_conditional_channel.py@d3b86dcdprivate repo diagnostic_v3/FINDINGS.md — five mechanisms measured to FAIL at closing the masked-fault gap (persistence, k-of-n items, block aggregation, model capacity, pooled multivariate T-squared), recorded so they are not retried. Protocol v2 is committed alongside as a MISS at 1/500 false alarms.
🔒 Sourcecubie-tf/docs/audit/tep_retest_2026-08/diagnostic_v3/FINDINGS.md@d3b86dcdprivate repo latency_masked_trio.json — component B's measured detection latency on IDV-3, IDV-9 and IDV-15, derived 2026-08-24. The delay is identically 359 samples (17.95 h) on all 500 held-out runs of each of the three faults, which is exactly the 360-sample window floor: every run alarms in its FIRST post-onset window, so the statistic contributes no latency of its own and the window length is the only lever on it. Because a constant delay is indistinguishable from a degenerate computation, the artifact also records the first window's own margin — minimum across runs x1.748 (IDV-3), x2.037 (IDV-9), x4.391 (IDV-15), with every post-onset window of every run crossing — so the floor is demonstrably reached with headroom rather than marginally. Produced read-only from the hash-pinned thresholds by tools/tep_masked_latency.py; it is the same category as the paper's figure generator and NOT a further evaluation of the held-out splits, so the exposure counts recorded in this claim are unchanged by it.
🔒 Sourcecubie-tf/docs/audit/tep_retest_2026-08/protocol_v3/latency_masked_trio.json@d3b86dcdprivate repo Paper — 'Detecting Controller-Masked Faults in the Tennessee Eastman Process at Zero False Alarms', 13 pages. PDF SHA-256 f30ae8f411e62131684435022795c46b265658783327578b76f5f435201b88cc. The LLM section is now built on a positive control and carries a like-for-like head-to-head against gpt-5 on seven held-out runs; the 11-page revision at c0a4236c (PDF f0525769…) carries the superseded blind-window probe, and the 10-page revision at 7b8bbd6e5 predates component B's measured latency.
🔒 Sourcecubie-research/paper/tep_masked_fault_detection.tex@0ec3818dprivate repo Head-to-head record — union detector vs gpt-5 on seven runs extracted from the held-out Rieth-2017 testing tables, both systems scored on byte-identical files. Union 7/7 correct; gpt-5 4/7. They agree on both fault-free runs (0/2 false alarms each) and on both unmasked positive controls IDV-1 and IDV-6 (2/2 each); the entire difference is the masked trio, 3/3 versus 0/3. Carries the sha256 of every input run, the extraction commands, and an explicit list of what the comparison does NOT establish — notably that seven runs is a demonstration, not a rate.
🔒 Sourcecubie-research/docs/evidence/2026-08-24-tep-llm-headtohead.md@0ec3818dprivate repo tep_head_to_head.py — scores the frozen union detector one run at a time so it can be placed beside a per-run LLM result. It imports cubie-tf's canonical tools/tep_conditional_channel.py and calls that module's own read_csv_table/residual/rolling against artifact sha256 0e88662d67b8d459552be5bf160d9cfa40c7648635fd90afc436e9a99eed0f13; the detector is not reimplemented and nothing in cubie-tf is modified. Component A is cited rather than re-executed, which the file states and justifies: its published run-level FDR on these faults is exactly 1.0000 or exactly 0.0000 and its fault-free alarm count is 0 of 500, all degenerate, so each single run's outcome follows with certainty.
🔒 Sourcecubie-research/tools/tep_head_to_head.py@0ec3818dprivate repo REPRODUCING_TEP_MASKED_FAULTS.md — end-to-end rebuild of component B from the public Harvard Dataverse corpus, with digests for every intermediate artifact and an explicit 'what reproduction does and does not establish' section.
🔒 Sourcecubie-research/paper/REPRODUCING_TEP_MASKED_FAULTS.md@c0a4236cprivate repo Corrected 50-active-cell mixed-scale schema-v3 reference, May recovery, grouped-Training evidence, and triple-kernel certification (CUB-3240, CUB-3251..3256). The certification explicitly denies confirmation and production eligibility. SUPERSEDED 2026-08-23 for current held-out status; retained as historical provenance.
🔒 Sourcecubie-research/docs/evidence/2026-08-01-tep-migration-certification.md@bd51e5e0private repo Public source pack — cubie-tep-source.zip (git-archive of the TEP detector crate: src/cusum.rs Page-CUSUM, two-pass calibration, OR-gate, src/bin/tep_detect.rs + tep_layout_search.rs). Enables source inspection without repo access; end-to-end build/run requires placing the pack in a matching cubie-tf workspace checkout. Bytes and regeneration notes are tracked in docs/resource-library/2026-07-04-source-packs.md.
Locked IPemail-gated IPsha256 fbdd757730dbcd9f883849ac79b7a9a67601ed0cee574d9800b810f88d140394 cubie-tep crate lib.rs — TEP fault detector crate root (STICKER layer over cubie-core; IDV-3/9/15 closed-loop-masked trio; Q16.16 f64-firewall; all algorithms executable, no stubs)
🔒 Sourcecubie-tf/cubie-tep/src/lib.rs@88ef7e83private repo cubie-tep/src/cusum.rs — CUB-1921 Page(1954) CUSUM aggregator: s_t=max(0,s_{t-1}+(|z|-k)); calibrate_two_pass (PASS A k[c]=E|z|+0.25 sigma, PASS B h[c]=1.5 x observed d00 peak => FAR=0 by construction); OR-gate any_fired()
🔒 Sourcecubie-tf/cubie-tep/src/cusum.rs@88ef7e83private repo docs/audit/empirical_peak.md — authoritative peak record: 100.00/100.00/100.00 @ d00 FAR=0.000% (0/960), commit 721394f, CUB-1921 CUSUM OR-gate; reproduced seeds {2024,5,73}; explicit in-sample calibration-identity scope disclosure; PCA beat factors 16.67x/33.33x/10.00x; full history table
🔒 Sourcecubie-tf/docs/audit/empirical_peak.md@88ef7e83private repo docs/audit/tep_recovery_status.md — closure of the 'perfect FAR=0% + three perfect FDRs' debt items; 100/100/100 across 3 V3-derived layouts; Pareto-improvement over prior 717e689 peak (41.12/31.62/29.12 @ FAR=0%)
🔒 Sourcecubie-tf/docs/audit/tep_recovery_status.md@88ef7e83private repo docs/audit/tep_oos_results.md — HONEST OOS boundary: status VERIFIED_BAD; default-config held-out proxy on d03/d09/d15 reports 0.00% FDR proxy; states the 100/100/100 headline 'remains an in-sample CUSUM calibration identity, not a held-out product claim'
🔒 Sourcecubie-tf/docs/audit/tep_oos_results.md@88ef7e83private repo cubie-tep/src/bin/tep_detect.rs — CLI detector over Rieth-2017 CSV schema (55 cols; fault inject sample 161); emits per-fault FDR proxy, FAR proxy, p99 detection latency
🔒 Sourcecubie-tf/cubie-tep/src/bin/tep_detect.rs@88ef7e83private repo cubie-tep/src/bin/tep_layout_search.rs — layout-search harness that reproduces the peak (--max-iters, --seed, --cusum default-on, --algorithm pareto)
🔒 Sourcecubie-tf/cubie-tep/src/bin/tep_layout_search.rs@88ef7e83private repo data/tep/layouts/best_found_by_search.json — search-discovered sticker permutation (seed 1729) used for the peak
🔒 Sourcecubie-tf/data/tep/layouts/best_found_by_search.json@88ef7e83private repo data/tep/layouts/_perfect_3wins_seed2024_cusum.json — one of the three reproduced perfect-3 layout artifacts (seed 2024)
🔒 Sourcecubie-tf/data/tep/layouts/_perfect_3wins_seed2024_cusum.json@88ef7e83private repo data/tep/README.md — dataset provenance (Russell/Chiang/Braatz d00/d03/d09/d15 from camaramm/tennessee-eastman-profBraatz; .dat files gitignored); local reproduction recipe. NOTE: the --use-braatz-baseline path alone underperforms PCA; the 100/100/100 peak requires the search layout + CUB-1921 CUSUM
🔒 Sourcecubie-tf/data/tep/README.md@88ef7e83private repo proofs/coq/tep/CUB_1952_CalibrationIdentity.v — Coq proof that FAR=0 on the calibration set is a definitional identity of h=alpha*max_t s_t with alpha>1
🔒 Sourcecubie-tf/proofs/coq/tep/CUB_1952_CalibrationIdentity.v@88ef7e83private repo proofs/lean/tep/CUB_1952_CalibrationIdentity.lean — Lean4 kernel of the CUB-1952 calibration identity
🔒 Sourcecubie-tf/proofs/lean/tep/CUB_1952_CalibrationIdentity.lean@88ef7e83private repo proofs/verus/tep/CUB_1952_calibration_identity_spec.rs — Verus spec of the CUB-1952 calibration identity
🔒 Sourcecubie-tf/proofs/verus/tep/CUB_1952_calibration_identity_spec.rs@88ef7e83private repo proofs/coq/tep/CUB_1953_OrGateCompleteness.v — Coq proof: OR-gate completeness (CUSUM OR binomial catches both fast and slow regimes)
🔒 Sourcecubie-tf/proofs/coq/tep/CUB_1953_OrGateCompleteness.v@88ef7e83private repo proofs/coq/tep/CUB_1954_V3LayoutDominance.v — Coq proof: V3/search layout dominance
🔒 Sourcecubie-tf/proofs/coq/tep/CUB_1954_V3LayoutDominance.v@88ef7e83private repo coq/CUB_1832_TepBelnapBaselineThreshold.v + coq/CubieCusumAggregator.v — the binomial-bounce (CUB-1832) and CUSUM aggregator kernels the OR-gate combines
🔒 Sourcecubie-tf/coq/CubieCusumAggregator.v@88ef7e83private repo TEP work evidence pack (library) — publishes the raw Rieth-2017 RData files (>1GB compressed) that are gitignored in the repo, plus cubie-tep crate, layouts, baseline stats, and benchmark/Opus-4.8 evidence. Manifest key cubie-tf/tep-work/2026-07-03/MANIFEST.md
Reproduce this result
# ===== HELD-OUT RESULT (protocol v3, 2026-08-23) — reproduce component B end to end =====
# Full step-by-step with digests at every stage: cubie-research paper/REPRODUCING_TEP_MASKED_FAULTS.md
# 0. Requirements: Python 3.14 + numpy/pandas/pyreadr, ~10 GB free disk, ~4 GB RAM, ~25 min. No GPU, no network after step 1.
python -m pip install numpy pandas pyreadr
# 1. Download all four .RData files from Harvard Dataverse doi:10.7910/DVN/6C3JR1 into data/tep/rieth2017/ and verify:
md5sum data/tep/rieth2017/*.RData # Testing digests are pinned in tools/tep_rieth_testing_to_csv.py (ARTIFACTS), which fails closed on drift
# 2. Convert the two TESTING files to canonical CSV (row order is load-bearing; use the converter, not an ad-hoc export):
python tools/tep_rieth_testing_to_csv.py --source data/tep/rieth2017/TEP_FaultFree_Testing.RData --kind fault_free_testing --output data/tep/rieth_retest/rieth_test_ff.csv --receipt data/tep/rieth_retest/receipt_ff.json
python tools/tep_rieth_testing_to_csv.py --source data/tep/rieth2017/TEP_Faulty_Testing.RData --kind faulty_testing --output data/tep/rieth_retest/rieth_test_faulty.csv --receipt data/tep/rieth_retest/receipt_faulty.json
sha256sum data/tep/rieth_retest/rieth_test_ff.csv # expect 605d62504c33404ef12c10a04e6a538d8fd11b9cb88cbb2780d7f2c1cb1509a2
# 3. Fit the residue channel from TRAINING ONLY (fit refuses any input whose basename contains 'Testing'):
python tools/tep_conditional_channel.py fit --root data/tep/rieth2017 --select-faults 3,9,15 --out docs/audit/tep_retest_2026-08/protocol_v3/residue_channel.npz
sha256sum docs/audit/tep_retest_2026-08/protocol_v3/residue_channel.npz # MUST equal 0e88662d67b8d459552be5bf160d9cfa40c7648635fd90afc436e9a99eed0f13, else you have not reproduced the evaluated detector — stop and reconcile before scoring
# 4. Score both held-out splits:
python tools/tep_conditional_channel.py score --artifact docs/audit/tep_retest_2026-08/protocol_v3/residue_channel.npz --input data/tep/rieth_retest/rieth_test_ff.csv --out results_ff.json --samples 960 --onset 161 # expect alarmedRunsAnyWindow = 0 of 500, maxRunMarginAnyWindow = 0.43759207958620044
python tools/tep_conditional_channel.py score --artifact docs/audit/tep_retest_2026-08/protocol_v3/residue_channel.npz --input data/tep/rieth_retest/rieth_test_faulty.csv --out results_faulty.json --samples 960 --onset 161 --faulty # expect IDV-3 1.0000 x1.982, IDV-9 1.0000 x4.166, IDV-15 1.0000 x5.537, IDV-4 0.0020 (carried by component A)
# 3b. OPTIONAL, read-only: re-derive component B's detection latency from the same frozen artifact. Reads no threshold it does not already carry and cross-checks every margin against results_faulty.json above, aborting on disagreement; it is not a further evaluation and does not change the exposure counts.
python tools/tep_masked_latency.py --artifact docs/audit/tep_retest_2026-08/protocol_v3/residue_channel.npz --input data/tep/rieth_retest/rieth_test_faulty.csv --result docs/audit/tep_retest_2026-08/protocol_v3/results_faulty.json --out docs/audit/tep_retest_2026-08/protocol_v3/latency_masked_trio.json # expect the 359-sample (17.95 h) window floor on 500/500 runs of each of IDV-3, IDV-9, IDV-15, with first-post-onset-window margins of at least x1.748, x2.037 and x4.391 respectively
# 5. Confirm the partition the union rule depends on, read off component A's PREVIOUSLY committed result:
python -c "import json; r=json.load(open('docs/audit/tep_retest_2026-08/result_faulty_seed73.json')); print('A at 1.000:',[f['fault_id'] for f in r['per_fault'] if f['run_fdr']==1.0]); print('A censored:',[f['fault_id'] for f in r['per_fault'] if f['run_fdr']==0.0])" # expect censored = [3, 9, 15]
# 6. Rebuild the paper and re-verify every figure against the committed record (aborts on any disagreement):
python paper/make_rieth_figures.py --tf /path/to/cubie-tf && cd paper && pdflatex tep_masked_fault_detection.tex && pdflatex tep_masked_fault_detection.tex
# ===== SUPERSEDED HISTORICAL MATERIAL BELOW — in-sample v8.1 replay, LLM comparison, legacy 100/100/100 =====
# Public source inspection: download the detector crate source pack.
curl -sL -o cubie-tep-source.zip 'https://lib.trustfortress.ai/objects/cubie-tf%2Fsource-packs%2F2026-07-04%2Fcubie-tep-source.zip' # hashes in https://lib.trustfortress.ai/objects/cubie-tf%2Fsource-packs%2F2026-07-04%2FMANIFEST.md
# This partial crate pack is not standalone-buildable: unzip into a matching cubie-tf checkout to build (`cargo run -p cubie-tep --bin tep_detect`), or read cubie-tep/src/cusum.rs directly for the CUSUM + calibration-identity logic.
# 1. Get the dataset (gitignored in repo; published in the library tep-work pack). Repo path: data/tep/README.md documents source camaramm/tennessee-eastman-profBraatz (d00/d00_te/d03_te/d09_te/d15_te .dat).
git clone https://github.com/iamdatanick/cubie-tf.git && cd cubie-tf && git checkout 88ef7e8303ed82b4e0af398960be172f8bb560db
# 2. Build the detector + search harness (release):
cargo build --release --package cubie-tep --bin tep_detect --bin tep_layout_search
# 3. Reproduce the search-discovered peak layout (seed 1729, 100 iters, CUSUM default-on):
python3 tools/tep_layout_search.py --max-iters 100 --seed 1729 # emits data/tep/layouts/best_found_by_search.json
# (Rust harness equivalent: ./target/release/tep_layout_search --seed 1729 --max-iters 100)
# 4. Run the detector per fault (Rieth convention: fault injected at sample 161):
./target/release/tep_detect --csv data/tep/d00_test.csv --summary-only # expect d00 FAR = 0.000% (0/960)
./target/release/tep_detect --csv data/tep/d03_test.csv --fault-inject-sample 161 --summary-only # expect IDV-3 FDR = 100.00%
./target/release/tep_detect --csv data/tep/d09_test.csv --fault-inject-sample 161 --summary-only # expect IDV-9 FDR = 100.00% (canonical window, samples 161+)
./target/release/tep_detect --csv data/tep/d15_test.csv --fault-inject-sample 161 --summary-only # expect IDV-15 FDR = 100.00%
# 5. Reproduce 60/60 saturation (whitepaper): run across 20 seeds x 3 algorithms (greedy/pareto/simanneal) at 1000-iter budget via tep_layout_search --algorithm {greedy,pareto,simanneal}.
# 6. Verify triple-kernel proofs of the calibration identity (per docker/ toolchains): coqchk on proofs/coq/tep/CUB_1952_CalibrationIdentity.v ; lean on proofs/lean/tep/CUB_1952_CalibrationIdentity.lean ; verus --crate-type lib on proofs/verus/tep/CUB_1952_calibration_identity_spec.rs.
# 7. (Honest boundary) OOS check: python tools/tep_oos_validation.py --dataset-dir data/tep --out docs/audit/tep_oos_results.md # default config returns 0.00% FDR proxy => confirms in-sample-only scope.
# 8. Download the merged v8.1 research evidence, supply your own key only in the current process, and treat the output as a new provider receipt.
curl -sL -o tep-v81-evidence.zip 'https://lib.trustfortress.ai/objects/cubie-research%2Ftep-v8.1-evidence%2F2026-08-03%2Fcubie-research-tep-evidence-bd51e5e.zip'
# PowerShell: Expand-Archive ./tep-v81-evidence.zip ./tep-v81-evidence; Set-Location ./tep-v81-evidence/cubie-research-bd51e5e/evidence/tep/may-2026-opus47-opus48
# Validate the exact 52-channel/960-row schedule without spending tokens: pwsh ./replay_opus_windows.ps1 -DatasetDir C:\path\to\tep-csv -Models claude-opus-4-7,claude-opus-4-8,claude-fable-5 -MeasurementCount 52 -FullCoverage -DryRun -OutFile ./dry-run.json
# Strict run: $env:ANTHROPIC_API_KEY='<your own key>'; pwsh ./replay_opus_windows.ps1 -DatasetDir C:\path\to\tep-csv -Models claude-opus-4-7,claude-opus-4-8,claude-fable-5 -MeasurementCount 52 -FullCoverage -OutFile ./full52-strict.json
# Adaptive-summary run: pwsh ./replay_opus_windows.ps1 -DatasetDir C:\path\to\tep-csv -Models claude-opus-4-7,claude-opus-4-8,claude-fable-5 -MeasurementCount 52 -FullCoverage -AdaptiveThinking -ThinkingEffort high -OutFile ./full52-adaptive.json; Remove-Item Env:ANTHROPIC_API_KEY
Verification checks
- Held-out fault-free scoring, re-measured from the committed raw JSON docs/audit/tep_retest_2026-08/protocol_v3/results_ff.json: artifactSha256 = 0e88662d67b8d459552be5bf160d9cfa40c7648635fd90afc436e9a99eed0f13, inputSha256 = 605d62504c33404ef12c10a04e6a538d8fd11b9cb88cbb2780d7f2c1cb1509a2, samplesPerRun 960, onsetSample 161, window 360, stride 8, runs 500, windowsPerRun 74, alarmedRunsAnyWindow = 0, maxRunMarginAnyWindow = 0.43759207958620044. Thresholds [0.4965034550825863, 3.957963428710814, 0.2013490512921977] on items [[20, xmeas_21, abs_mean_z], [16, xmeas_17, var_z], [20, xmeas_21, acf_pow6]].
- Held-out faulty scoring, re-measured from docs/audit/tep_retest_2026-08/protocol_v3/results_faulty.json: 20 faults x 500 runs. Component B alone reaches detectionRatePostOnset = 1.0000 on 19 of 20 faults and 0.0020 on IDV-4. Worst-run post-onset margins on the masked trio: IDV-3 x1.982, IDV-9 x4.166, IDV-15 x5.537. Component A covers IDV-4 at 1.0000, so the union is 20/20.
- Threshold independence, which is what bounds the union false alarms before scoring: cmd_fit derives each threshold per item and each fault's item per fault, so the three retained thresholds are numerically identical to those committed under the earlier eighteen-item variant. The three-item channel's alarm set is therefore a strict subset of the eighteen-item channel's.
- Train/test separation is enforced in code, not by discipline: tep_conditional_channel.py fit refuses any input whose basename contains 'Testing', so no testing byte can reach the coefficients, the item selection or the thresholds. Training is split three ways — coefficients from clean runs 1-250, item selection on clean 251-375 plus faulty 1-250, thresholds from clean 251-500.
- Paper artifacts at cubie-research c0a4236c5eba4dd660378ba64f0ab110417b93fb: tep_masked_fault_detection.pdf SHA-256 = f052576958bf86953516b2e768b6df331ecb42ed107a0ef293fa37c154f60619 (11 pages), tep_masked_fault_detection.tex SHA-256 = bb86205b23a7ed7e0bab4030476c12fb8dbf85bf65e1a8f416662d70e5fa3e46, REPRODUCING_TEP_MASKED_FAULTS.md SHA-256 = 27f362599ca71543e914883e3df2763bb73b58266b8c719aee7e4c630b322970. The figure generator make_rieth_figures.py cross-checks nine derived quantities against the committed result JSONs and aborts rather than typesetting on any disagreement; all nine pass.
- Latency artifact at cubie-tf d3b86dcddd8927dc38ed6c8dc462f645404bf409: tools/tep_masked_latency.py SHA-256 = 755dfeb4dfa38e54224275166bf8e818b35bbc30048246f6c15f0c158192931f, latency_masked_trio.json SHA-256 = df158190975f8183b8bbf9501ea1e0ddeb5a7f058e7e119010523c62c8124823. The tool re-derives each fault's worst-run margin and alarmed-run count from the frozen thresholds and aborts unless both agree with the committed results_faulty.json; all six checks pass (x1.982320597, x4.166148755, x5.536743 and 500/500 alarmed on each of IDV-3, IDV-9, IDV-15).
- V8.1 whitepaper PDF SHA-256 = 80c017a2f9d8d6142fd964d6e47964d9d0973e1634d181014c78fb22bc8a7e75 (9 pages, merged at cubie-research bd51e5e). SUPERSEDED for current held-out status; retained as historical provenance.
- Complete seed-19 continuous replay: 3,840 rows; 2,400/2,400 fault-active alarms; 125/1,440 clean-labelled alarms, all pre-onset; d00 0/960; 3,715/3,840 correct (96.745%).
- Common full-coverage schedule: 52 channels, 48 non-overlapping windows per file, 192 decisions per system. Cubie 184/192; strict Opus 4.7 96/192, Opus 4.8 85/192, Fable 5 105/192; adaptive Opus 4.7 102/192, Opus 4.8 86/192, Fable 5 102/192.
- Historical V6 whitepaper PDF SHA-256 = 8569EC5FF26A77E3B55591E11FF8A75DB26EBEB82B523E94F376696AC60B9B05 (non-watermarked edition; retained as historical provenance).
- Whitepaper body (extracted): 'the final detector achieves FDR_IDV-3 = 100.00%, FDR_IDV-9 = 100.00%, FDR_IDV-15 = 100.00% at d00 FAR = 0.000%' and 'the zero-FAR boundary is not statistical inference but a calibration identity' and 'All 20x3 = 60 runs converge to the saturation peak ... 60/60 perfect-3'.
- empirical_peak.md history table row 2026-05-26 commit 721394f: 100.00%/100.00%/100.00%, 3 wins, d00 FAR=0.000% (0/960 by h-calibration construction).
- cusum.rs calibrate_two_pass: PASS B computes h[c] = (peaks[c] as f64 * 1.5) => guarantees P(fire on d00)=0 (calibration identity, matches CUB-1952).
- All 18 cited repo paths confirmed at evidence commit 88ef7e8303ed82b4e0af398960be172f8bb560db via gh api repos/iamdatanick/cubie-tf/contents/<path>?ref=88ef7e8303ed82b4e0af398960be172f8bb560db (each returned a blob sha).
- Triple-kernel suite CUB-1950..CUB-1959 present in proofs/{coq,lean,verus}/tep/ (git ls-files); CUB-1952/1953/1954 are the calibration-identity / OR-gate-completeness / layout-dominance proofs.
- OOS honesty gate: docs/audit/tep_oos_results.md status VERIFIED_BAD, 0.00% FDR proxy at default config, explicitly labels the headline 'in-sample CUSUM calibration identity, not a held-out product claim'.
Known boundaries / open gaps
- CLOSED 2026-08-23 for the union detector, still OPEN for the legacy CUSUM/layout configuration. The clean transfer to the Rieth-2017 testing splits (Harvard Dataverse DOI 10.7910/DVN/6C3JR1) HAS now been run: 0/500 fault-free runs alarmed, 20/20 faults at run-level 1.0000, thresholds frozen and hashed before scoring. What remains open is (a) the legacy default-config detector that docs/audit/tep_oos_results.md records as VERIFIED_BAD at 0.00% FDR proxy — that document describes a different configuration and was NOT re-run, (b) an untouched first-exposure confirmation set, since the fault-free testing file has now been scored three times, and (c) Q16.16 fixed-point integration of component B into the shipped Rust detector, which was not exercised. The 100/100/100 @ FAR=0 result is in-sample (calibration identity), and the whitepaper + audit docs both scope it that way. Checked: whitepaper (library), cubie-tf audit docs + code, data/tep. Not a missing-evidence gap — it is an explicit scope boundary of the claim.
- The raw Braatz/Rieth .dat/.RData datasets are gitignored in cubie-tf (only committed baseline_stats.json / layouts / MANIFEST.sha256). They ARE published in the library tep-work evidence pack (>1GB compressed, cubie-tf/tep-work/2026-07-03/), but that manifest does not enumerate per-file object keys, so exact per-dataset object URLs could not be captured without extracting the pack.
- Repo attribution nuance: the whitepaper author line reads github.com/iamdatanick/cubie-math and is dated 2026-05-26 (pre-monorepo). The live, current code + proofs are in iamdatanick/cubie-tf (cubie-math was frozen and absorbed into the monorepo on 2026-05-27). All blob URLs above point to the live cubie-tf/main location.
- The corrected 50-active-cell mixed-scale schema-v3 reference and v8.1 evidence live in cubie-research at immutable merge commit bd51e5e. They contain reference semantics, offline tools, exact May source/control recovery, CUB-3240/CUB-3251..3256 proof triples, a corrected Training-only baseline, grouped-CV specialist evidence, and full Cubie/Anthropic replay receipts. Fault 3 reached 92.6% and fault 15 reached 29.6%, while fault 9 reached 0.8% and failed its predeclared 1% floor; the partial profile is therefore not confirmation-eligible. No sealed corrected provider bundle, untouched confirmation replay, or production authority exists. The legacy 100/100/100 result remains attributable only to its historical cubie-tf snapshot. SUPERSEDED 2026-08-23: the fault-9 floor failure recorded here no longer describes the current evidence — under the protocol-v3 union, fault 9 is detected on 500/500 held-out runs at a worst-run margin of x4.166. The grouped-Training figures in this paragraph are retained as historical provenance for the 50-active-cell profile, which is a different configuration.
- For 0 alarmed runs out of 500 representative independent runs, the exact one-sided 95% binomial upper bound is about 0.5974% per run. The supported wording is observed 0/500, not FAR = 0.
- PIN FRESHNESS: the held-out evidence below is pinned to the head commits of two OPEN pull requests — cubie-tf#1398 at d3b86dcddd8927dc38ed6c8dc462f645404bf409 and cubie-research#49 at c0a4236c5eba4dd660378ba64f0ab110417b93fb. Those SHAs are immutable and reachable now, but if either PR is squash-merged the canonical history will carry different SHAs and this claim must be re-pinned to the merge commits. Until both merge, the evidence lives outside the repositories' main branches. RE-PINNED ONCE ALREADY, on 2026-08-24, when the latency derivation moved both branch heads (cubie-tf#1398 from 4a8136d83, cubie-research#49 from 7b8bbd6e5); the superseded SHAs remain valid immutable commits but point at the pre-latency paper, which still asserts that no detection-delay figure exists. Cite the heads named above, not the earlier pair.
- Single-item fragility of the masked trio: each of IDV-3, IDV-9 and IDV-15 rests on ONE window statistic, and two of the three on the same cell (xmeas_21, reactor cooling water outlet temperature). A sensor fault or a re-tuned controller affecting xmeas_21 would remove two of the three detections at once. Masking is a property of the controller, so a different control scheme would redistribute which faults are masked and require re-selection — though not a change of method. All results are for the single decentralised control scheme distributed with the Rieth corpus.
- CLOSED 2026-08-24. This gap previously read that component B reports run-level detection only and that there is NO detection-delay figure for IDV-3, IDV-9 or IDV-15. There now is one. Component B's latency is bounded below by its own 360-sample rolling window (about 18 hours at the 3-minute sampling interval): the first window lying wholly after onset starts at sample 161 and cannot declare before its last sample at 520, a floor of 359 samples or 17.95 hours. The held-out data gives that floor exactly — on all 500 runs of each of the three faults the first post-onset window already crosses, so the measured delay is identically 359 samples with no spread whatever, and the window length is the only lever on it. Because a delay that is constant across 1500 runs is indistinguishable from a degenerate computation, the first window's own margin is reported with it: minimum across runs x1.748 (IDV-3), x2.037 (IDV-9), x4.391 (IDV-15), every post-onset window of every run crossing. STILL IN FORCE, and unchanged by the closure: component A's much shorter latency numbers must NOT be quoted for a fault only component B detects, nor pooled with B's into a single latency figure for the union — the two components measure on different timescales by construction. The derivation is read-only with respect to the evaluation (frozen hash-pinned thresholds, nothing can feed back into a detector parameter), the same category as the paper's figure generator, so the exposure counts recorded elsewhere in this claim are NOT incremented by it.
- Five mechanisms were measured and FAILED to close the gap on training data — persistence over consecutive windows, requiring k distinct items, block aggregation, higher model capacity, and a pooled multivariate T-squared over all 208 moments. They are recorded in docs/audit/tep_retest_2026-08/diagnostic_v3/FINDINGS.md so they are not retried, and protocol v2 is committed as a MISS (1/500 false alarms) rather than deleted.