Tenstorrent Blackhole: Performance Claims, Open Stack, and Due-Diligence Risks
Vendor benchmarks, PyTorch and Hugging Face compatibility, and the diligence checklist AGICY applies before Phase 1 fleet scale-up.
Tenstorrent reports ~350 tok/s on DeepSeek-R1-class inference at roughly ~$6/M tokens — compared to ~$30/M on GB300-class systems in vendor presentations. The open stack (RISC-V, PyTorch, Hugging Face) is compelling for sovereign campuses, but SDK maturity, yield, firmware stability, and leadership continuity remain standard pre-production risks requiring independent validation.
Key Takeaways
- DeepSeek-R1 benchmark:~350 tok/s on Galaxy-class hardware (vendor-claimed); ~$6/M tokens vs ~$30/M on GB300-class (vendor comparison).VENDOR-CLAIMED
- Open ISA: RISC-V enables firmware and microcode audit for EU regulated workloads.VERIFIED
- Framework compatibility: PyTorch and Hugging Face model import claimed for mainstream open-weight models.VENDOR-CLAIMED
- SDK maturity risk: TT-Metalium / TT-Forge toolchain still maturing vs decade-old CUDA ecosystem.DESIGN TARGET
- Yield & firmware: New process nodes and multi-chip Galaxy integration require burn-in and OTA stability programmes.DESIGN TARGET
- Leadership continuity: Emerging silicon vendors face key-person and roadmap execution risk — standard vendor DD.DESIGN TARGET
- AGICY contingency:Phase 1 delay triggers interim AMD MI300X air-cooled racks per vendor-neutral policy.VERIFIED
Performance Claims: DeepSeek-R1 and $/Token Economics
Tenstorrent's public Galaxy demonstrations highlight DeepSeek-R1 inference — a reasoning-heavy open model relevant to education, research, and public-sector analytics. Vendor materials cite approximately 350 tokens per second sustained decode on Galaxy configurations, with implied economics near $6 per million output tokens.
The comparison baseline in vendor decks — GB300-class GPU systems at roughly $30/M tokens — is illustrative. AGICY treats all vendor-side comparisons as hypotheses until reproduced on identical batch sizes, context lengths, and power meters in Vasilikos acceptance testing.
~$6/M output tokens vs ~$30/M on GB300-class systems in Tenstorrent marketing comparisons. AGICY Phase 1 fleet model uses 300 tok/s sustained fleet-average for capacity planning — conservative vs peak vendor demos.VENDOR-CLAIMED
Open Stack: RISC-V, PyTorch, Hugging Face
Sovereign buyers care as much about software portability as raw speed. Tenstorrent ships TT-Metalium and TT-Forge alongside PyTorch integration paths, with stated compatibility for Hugging Face model hubs — reducing the friction of deploying Llama, Mistral, and DeepSeek-class weights on EU soil.
- RISC-V: Open ISA supports third-party security review — aligned with AGICY auditability gate.
- PyTorch-first: Matches AGICY training/serving policy — no mandatory proprietary graph compiler.
- Model agnosticism: Client BYOM supported; AGICY does not require exclusive model hosting.
Maturity gap remains: CUDA's library depth (cuDNN, NCCL, TensorRT) accumulated over fifteen years. TT-Forge must prove parity on your models — not benchmark charts alone.
Due-Diligence Risk Register
| Risk | Description | AGICY Mitigation |
|---|---|---|
| SDK maturity | Toolchain gaps vs CUDA for exotic ops and custom layers | Pre-COD model compatibility matrix; vLLM-class fallbacks |
| Yield / firmware | Multi-chip Galaxy integration and early-node defect rates | 50-unit spare pool; staged burn-in per deployment plan |
| Leadership continuity | Key-person dependency in emerging silicon vendors | Technology Committee quarterly review; multi-vendor slots |
| Benchmark inflation | Vendor peak vs fleet-sustained throughput divergence | 300 tok/s fleet-average in published basis; independent acceptance tests |
What Independent Validation Must Prove
Before AGICY publishes customer-facing SLAs on Galaxy capacity, the Technology Committee requires:
- Reproducible throughput on target models at 65% fleet utilisation (Phase 1 management case).
- Power draw verification at 9 kW unit TDP × 1.15 PUE design point.
- Failover behaviour across Ethernet scale-out under single-rack loss.
- TT-Forge compile success rate for top-20 client model architectures.
Until those gates pass, public headline capacity remains 17T tokens/year as a target derived from deterministic fleet math — not a vendor marketing claim.DESIGN TARGET
AGICY Relevance
AGICY selects Galaxy for Phase 1 because it passes preliminary four-gate review — but selection is conditional on ongoing diligence, not faith in keynote slides.
Civil workloads (healthcare triage, university LLM sandboxes, municipal digital services) benefit from open-weight models on auditable RISC-V hardware — provided SDK and firmware risks are actively managed.
Per AGICY's vendor-neutral policy, shipment delay triggers interim AMD MI300X racks; superior TCO from any vendor triggers integration — AGICY never becomes captive to Tenstorrent roadmap timing.
Request Diligence Briefing
Institutional buyers can access AGICY's Technology Committee summary, fleet acceptance criteria, and vendor-neutral contingency plan under NDA.
Contact Investor RelationsReferences & Primary Sources
- Tenstorrent — Galaxy GA announcements and DeepSeek-R1 demo materials (2026).
- AGICY Galaxy fleet specification — fleet-average 300 tok/s basis.
- AGICY vendor-neutral policy — stress-test and exit ramps.
- AGICY token-fleet basis note — 17T vs 38.5T public narrative.
- AGICY Research synthesis — Jim Keller / Blackhole video keypoints (Jul 2026).