TF Logs agentic-waste findings: GPU deadlock blast radius, VRAM lock, request-waste rate across ~45.3M real telemetry rows
Verified
Scope boundary: Re-measured from the committed result JSONs (not the prior memory dossier). ~99.5% output<input waste is scoped to the Azure-2024 trace (44.1M of the ~45.3M rows); only the 66.85% blast-radius figure appears in a whitepaper.
Claim
TrustFortress's "agentic waste" analysis over five real public datasets quantifies structural inefficiency in production LLM inference fleets. Re-measured from the committed analysis result files: (1) TP4 NVLink cascade deadlock freezes 66.85% of GPU fleet capacity during events (105,226 of 157,411 LLM GPU-intervals had SM duty==0; full_tp4_blast_radius.capacity_pct=66.8479, 140/141 containers); (2) 51.76% of all fleet VRAM is locked at any moment by the TP4 cascade (VRAM-GB-interval fraction (222,368.7+1,649,527.9)/3,616,626.1); (3) ~99.5% request waste = output<input, measured over the Azure-2024 trace only (43,872,851 of 44,107,694 requests = 99.47%). The total analyzed corpus re-summed from the result JSONs is ~45.29M distinct dataset rows (Alibaba 157,411 intervals + Azure-2023 28,185 + Azure-2024 44,107,694 + Azure-LMM-2025 1,000,000 + MaverIQ 1,684). SCOPING: Only the 66.85% figure appears in a whitepaper (V8 bare-metal .md, and its footnote marker [3] mis-points to the Rubik's-cube paper). The 51.76%, 99.5%, and row-total figures are backed by the JSON result files + INTC Evidence Pack, NOT by any whitepaper. The prior memory dossier's headline "44.9M rows / 99.5% of all 44M" is an approximation; the re-measured distinct-row total is ~45.3M, and 99.5% is scoped to Azure-2024 (44.1M of the 45.3M), not the whole corpus. The current live public pack at prefix cubie-tf/agentic-waste-evidence/2026-07-04 publishes the four headline JSONs needed for the 66.85% / 51.76% / ~99.5% reproduction; row-total support JSONs remain repo/private until the next owner-authorized publish run.
Direct downloads (public — lib.trustfortress.ai)
Public reproduction pack — REPRODUCE.md with copy-paste curl+verify+re-derive commands for the 66.85% / 51.76% / ~99.5% figures (no repo access required), plus the intc MANIFEST mapping each number to its JSON path.
intc-v1.0-tp4_nvlink_results.json — TP4 NVLink cascade: full_tp4_blast_radius {intervals:105226, capacity_pct:66.8479, containers:140}; vram_locked_by_tp4.pct_total_vram_locked:51.76; total_llm_intervals:157411. Source of the 66.85% and 51.76% headline numbers.
⬇ Downloadsha256 7d1a90bfc6036e7d6eb8483ba1ddb4bf69cc69003a306016368d221286119bc9 intc-v1.0-hunter_results.json — DCGM threshold sweep + ITL backpressure: exact_zero_sm {count:105226, pct:66.8479}; itl_proxy_code.gen1_pct:15.9979 (2,688,232 single-token aborts); retry chains full_chains_ge3:127; lmm context_spikes_10x_median:93157, max effective ctx 28,375,569 tokens.
⬇ Downloadsha256 6fb720db2600441e6a7155d4a2ac6ee8cdc023de65fceef016edc7742d979685 intc-v1.0-azure_2024_unconstrained.json — Azure-2024 request-waste: code total_rows:16,803,695 output_lt_input pct:99.2115; conv total_rows:27,303,999 output_lt_input pct:99.6252. Combined 43,872,851/44,107,694 = 99.47% ≈ the ~99.5% waste headline.
⬇ Downloadsha256 4a2a451e4e104bc543cce67aae42f8446513ea7160c7c32a3a6a880c2bfada06 intc-v1.0-alibaba_v2026_squatter_metrics.json — original GPU-squatter run (pre-TP4-reinterpretation): chronic_squatter_rows:7553, containers:131, fleet_capacity_destroyed_pct:4.7983, wasted_kwh:44.0592. Shows the 4.8% original filter that only caught GPU 0.
⬇ Downloadsha256 2e4adf2dcd79233d603e1c68498434a469413e292b8843bca7bc5131844d43ad
Source evidence (pinned)
intc-v1.0-MANIFEST.md — evidence-pack manifest. Lists per-file sha256 that MATCH the committed copies exactly (tp4=7d1a90bf…, hunter=6fb720db…, azure_2024=4a2a451e…, alibaba=2e4adf2d…); there is no upstream-vs-committed divergence. Maps each canonical dashboard number to its JSON path (66.85% -> full_tp4_blast_radius.capacity_pct, 51.76% -> vram_locked_by_tp4, 87.53% DCGM blind spot derived).
🔒 Sourcecubie-tf/demo/evidence/intc-v1.0/intc-v1.0-MANIFEST.md@88ef7e83private repo INTC_Evidence_Pack_v1.0.md — reproduce documentation: names analysis scripts scripts/tp4_nvlink_breakdown.py and hunter.py, the results/*.json outputs, and the causal chain (99.5% output<input -> KV-cache pin -> 66.85% fleet frozen). Confirms scripts live in the raw bundle, not the git repo.
🔒 Sourcecubie-tf/intel-submissions/cubie-intel-submission/intel-submission/evidence/INTC_Evidence_Pack_v1.0.md@88ef7e83private repo CUBIE_MASTER_WHITE_PAPER_V8_0_BARE_METAL.md — the ONLY whitepaper that states an agentic-waste number: 'freezing up to 66.85% of total data center fleet capacity without triggering hypervisor alerts [3]'. Caveat: footnote [3] mis-points to Rokicki 2014 (Rubik's cube diameter), a broken citation. Does NOT mention 51.76%, 99.5%, or the row total.
🔒 Sourcecubie-tf/docs/CUBIE_MASTER_WHITE_PAPER_V8_0_BARE_METAL.md@88ef7e83private repo Library raw-data pack (LIVE, HTTP 200) — datacenter-crash-logs/2026-07-02 MANIFEST with real whole-file SHA-256s. Raw bundle TF Logs 2 LARGE.zip (410,827,647 bytes, sha256 3141ab66a0ffe8f855aa503670d8664c194733cf36f438942e1230af09c5b7b6) split into part001 (sha 1a399c21...) + part002 (sha 01d7faa5...); plus tf-intc-analysis-pack-20260527-trimmed.zip (263,520 bytes, sha ecceb2a2...) containing the analysis scripts.
Library raw-data object part001 (LIVE, HTTP 200), 209,715,200 bytes — first half of the canonical TF Logs 2 LARGE evidence archive (Alibaba GenAI GPU telemetry, Azure LLM 2023/2024 + LMM 2025 traces, MaverIQ CSVs, vLLM OOM/DCGM artifacts).
Locked IPemail-gated IPsha256 1a399c21046176a99d1ddabfda9337e9bed4411dc704fc52159ef9c5b5771007
Reproduce this result
# Public reproduction (no repo access). Full guide: https://lib.trustfortress.ai/objects/cubie-tf%2Fagentic-waste-evidence%2F2026-07-04%2FREPRODUCE.md
curl -sL -o tp4.json 'https://lib.trustfortress.ai/objects/cubie-tf%2Fagentic-waste-evidence%2F2026-07-04%2Fintc-v1.0-tp4_nvlink_results.json'
curl -sL -o az2024.json 'https://lib.trustfortress.ai/objects/cubie-tf%2Fagentic-waste-evidence%2F2026-07-04%2Fintc-v1.0-azure_2024_unconstrained.json'
sha256sum tp4.json az2024.json # tp4 7d1a90bfc6036e7d6eb8483ba1ddb4bf69cc69003a306016368d221286119bc9 ; azure_2024 4a2a451e4e104bc543cce67aae42f8446513ea7160c7c32a3a6a880c2bfada06
Reproduce the 66.85% and 51.76% headlines directly from the JSON: python3 -c "import json;d=json.load(open('tp4.json'));print(d['full_tp4_blast_radius'])" -> EXPECT {'intervals':105226,'capacity_pct':66.8479,'containers':140,...} ; and d['vram_locked_by_tp4']['pct_total_vram_locked'] -> EXPECT 51.76 (=(222368.7+1649527.9)/3616626.1).
Reproduce the ~99.5% request-waste figure: fetch intc-v1.0-azure_2024_unconstrained.json, then python3 -c "import json;d=json.load(open('az2024.json'));o=d['code']['output_lt_input']['count']+d['conv']['output_lt_input']['count'];t=d['code']['total_rows']+d['conv']['total_rows'];print(o,t,round(100*o/t,4))" -> EXPECT 43872851 44107694 99.4676 (i.e. 99.5% output<input, scoped to Azure-2024).
Repo/private row-total self-check until the row-total support JSONs are uploaded: from demo/evidence/intc-v1.0, Alibaba 157411 + Azure2023 (8819+19366) + Azure2024 (16803695+27303999) + LMM2025 1000000 + MaverIQ 1684 = 45,294,974 distinct rows (~45.3M). tools/publish_reproduce_data.ps1 now includes the Azure-2023, LMM-2025, and MaverIQ support JSONs for the next public pack.
Regenerate the JSONs from raw data (full reproduce): download the raw bundle from the LIVE library pack — curl -O https://lib.trustfortress.ai/objects/cubie-tf%2Fdatacenter-crash-logs%2F2026-07-02%2Ftf-logs-2-large-20260412.zip.part001 (and .part002), reassemble ('copy /b name.part001+name.part002 name.zip'), verify sha256 == 3141ab66a0ffe8f855aa503670d8664c194733cf36f438942e1230af09c5b7b6 ; unzip; then run the analysis scripts (scripts/tp4_nvlink_breakdown.py -> results/tp4_nvlink_results.json ; hunter.py -> results/hunter_results.json) which ship inside tf-intc-analysis-pack-20260527-trimmed.zip (sha256 ecceb2a2e201b1a8871a4acb558717dcdafffe285a41c1a928ea7e618f221114) from the same library pack, per INTC_Evidence_Pack_v1.0.md.
Verification checks
- JSON headline metrics reproduce exactly: full_tp4_blast_radius.capacity_pct=66.8479 (105226/157411 zero-SM intervals), vram_locked_by_tp4.pct_total_vram_locked=51.76, Azure-2024 output<input = 43872851/44107694 = 99.4676%. Confirmed by base64-decoding the GH-API-served file content and re-summing.
- Whitepaper anchoring confirmed by grep: ONLY 66.85% appears (V8 bare-metal .md), and its footnote [3] mis-points to Rokicki 2014 (Rubik's cube). Greps for 51.76|99.5|45.3|44.9|VRAM.lock returned zero hits — those three are result-JSON-only.
- Library object availability confirmed by HTTP status: datacenter-crash-logs/2026-07-02 MANIFEST + part001 return 200 with the manifest's SHA-256 table; tf-logs-evidence/2026-07-04 MANIFEST returns 404 (object store empty — it is a locator doc, not a populated pack).
- Committed-copy hash self-consistency confirmed 2026-07-05: the four registry download sha256 values match the committed demo/evidence/intc-v1.0 JSON files and intc-v1.0-MANIFEST.md exactly (tp4=7d1a90bf..., hunter=6fb720db..., azure_2024=4a2a451e..., alibaba=2e4adf2d...).
- Live public-link HEAD verification on 2026-07-06 returned HTTP 200 for REPRODUCE.md, MANIFEST.md, and the four headline JSON objects under cubie-tf/agentic-waste-evidence/2026-07-04/. The Azure-2023, LMM-2025, and MaverIQ row-total support object URLs returned HTTP 404, so they are intentionally not listed as current public downloads.
Known boundaries / open gaps
- Analysis scripts (scripts/tp4_nvlink_breakdown.py, hunter.py) are NOT committed to the GitHub repo — they exist only inside the raw bundle (tf-intc-analysis-pack-20260527-trimmed.zip on the library, sha ecceb2a2..., per INTC_Evidence_Pack_v1.0.md). A reader can re-hash and re-read the precomputed JSONs from GitHub, but must download the library zip to regenerate them from raw data. Not unfindable (the script names + container zip are documented), but they lack a direct GitHub blob URL.
- The library pack referenced by the repo locator doc docs/resource-library/2026-07-04_tf-logs-evidence.md — prefix cubie-tf/tf-logs-evidence/2026-07-04/ — is a locator only: its object store is EMPTY (MANIFEST returns HTTP 404; api/list returns []). The live, populated raw-data pack is cubie-tf/datacenter-crash-logs/2026-07-02/ (HTTP 200). Referenced-but-not-yet-populated, not a true evidence gap. Checked: library api/list, api/search, direct /objects HEAD.
- Number-reconciliation note (accuracy, not a gap): the prior memory dossier states 44.9M rows and 43,666,866 Azure-2024 rows; the committed result JSONs give Azure-2024 = 44,107,694 and a 5-dataset distinct-row total of 45,294,974 (~45.3M). Registry uses the re-measured JSON figures. Also the dossier/whitepaper phrase '99.5% of all 44M' is scoped in the data to Azure-2024 (44.1M) specifically, not the whole 45.3M corpus.
- Current public-pack boundary: the four headline JSONs are live for no-repo reproduction, but the row-total support JSONs (Azure-2023, LMM-2025, MaverIQ) are not yet uploaded in the 2026-07-04 prefix. The publisher now includes them for the next authorized upload.
- Hash provenance is now aligned for the four committed result JSONs: cite the registry download rows or demo/evidence/intc-v1.0/intc-v1.0-MANIFEST.md for SHA-256 verification of the GitHub copies.