Cerebras CS-3: Wafer-Scale Training (July 2026)
Inside the WSE-3 wafer-scale engine — 4 trillion transistors, 900,000 AI cores, and why AGICY evaluates CS-3 as Phase 2 optionality (not an owned fleet).
Updated 19 August 2026. Cerebras announced CS-4 on 18 August 2026 (Nexus rack, three WSE-3 Turbo wafers — not a new WSE-4 die). This article is the 5 July 2026 CS-3 deep-dive. Living product page: /hardware/cerebras. AGICY has not ordered CS-3 or CS-4.
The Cerebras CS-3, powered by the WSE-3 wafer-scale engine, was Cerebras' shipping wafer-scale system when this article was published (5 July 2026). With 4 trillion transistors, 900,000 cores, and 44 GB of on-chip SRAM, Cerebras claims ~1,800 tokens/sec on Llama 3.1 70B. Unlike inference-only accelerators, the CS-3 supports training up to 24 trillion parameters via weight streaming (vendor-claimed). AGICY evaluates CS-3 / CS-4 as Phase 2 optionality — not an owned training hall.
WSE-3: The World's Largest AI Chip
Cerebras' third-generation Wafer Scale Engine occupies an entire 300 mm silicon wafer — 56× larger than the biggest discrete GPU. By keeping the die intact, data never leaves the chip boundary, eliminating the inter-chip communication overhead that plagues multi-GPU clusters.
- 4 trillion transistors— fabricated on TSMC 5 nm
- 900,000 AI-optimised cores — sparse-linear-algebra engines
- 44 GB on-chip SRAM — no external HBM required
- 21 PB/s memory bandwidth — 1,000× more than an H100
- Weight Streaming architecture — decouples compute from model storage, enabling linear scaling across 64+ CS-3 systems for models up to 24T parameters
Measured on a single CS-3 system via the Cerebras Inference API. This is approximately 7× the throughput of Groq LPU and 20× faster than NVIDIA H100 tensor-core serving for the same model at comparable batch sizes.
Cerebras vs Groq LPU: Speed Comparison
| Dimension | Cerebras CS-3 (WSE-3) | Groq LPU |
|---|---|---|
| Architecture | Wafer-scale, weight streaming | LPU (Language Processing Unit), SRAM-only |
| Llama 3.1 70B speed | ~1,800 tok/s | ~250 tok/s |
| Llama 3.1 8B speed | ~4,000 tok/s | ~800 tok/s |
| Training | ✓ Up to 24T parameters | ✗ Inference only |
| On-chip memory | 44 GB SRAM | 230 MB SRAM per chip |
| Max model size (single system) | 24T params (weight streaming) | ~70B params |
| Sovereign deployment | On-prem & private cloud | GroqCloud only (US-hosted) |
Business Milestones
Cerebras' transition from stealth hardware lab to public company has been one of the fastest in semiconductor history:
- NASDAQ IPO (CBRS) — May 2026:Listed at a $12 B valuation after clearing CFIUS national-security review, making it the first wafer-scale chip company on a public exchange.
- $20B OpenAI training contract: Multi-year deal to supply CS-3 clusters for frontier model pre-training, displacing a portion of NVIDIA H100/B200 capacity.
- AWS Marketplace partnership: CS-3 on-demand instances available in
us-east-1andeu-west-1, lowering the barrier for enterprises to benchmark wafer-scale performance. - 92% YoY revenue growth (Q1 2026): Driven primarily by sovereign-cloud contracts outside the US, signalling strong international demand for export-unrestricted AI training hardware.
Why AGICY Deploys Cerebras for Phase 2
AGICY's infrastructure roadmap is deliberately multi-silicon:
- Phase 1 — Tenstorrent RISC-V (inference): 1,801 Galaxy servers targeting 17T tok/year for low-latency serving of open-weight models. Campus is pre-construction (COD TARGET H2 2027) — not operational.
- Phase 2 — Cerebras CS-3 (training):Adds wafer-scale training capacity so EU customers can fine-tune and pre-train models up to 24T parameters without data leaving EU jurisdiction. Deployment target: Q1 2027.
- Phase 3 — Hybrid orchestration: Unified scheduler routes workloads to the optimal silicon — Tenstorrent for inference cost-efficiency, Cerebras for training throughput — across the Vasilikos campus.
The combination of RISC-V inference and wafer-scale training creates a vertically sovereign stack: no CUDA lock-in, no US-CLOUD-Act exposure, and no dependency on a single silicon vendor.
“The future of sovereign AI requires both inference and training capability on EU soil. Inference without training is renting intelligence — training on sovereign hardware is owning it.”
GPU leasing, rent, buy, or colocation
Search intent we actually fulfil on /leasing: GPU leasing, rent GPU server, buy AI accelerator (client-site finance under T&Cs), or AI colocation / GPU colo — also typed as collocation. Planned Cyprus campus; not a live hall.
This page is Cerebras CS-3 — Phase 2 quote-only evaluation. GPU leasing / rent / buy / colocation via /leasing. Pre-construction.
- Tenstorrent Galaxy — Phase 1 · list bands L1–L3
- Cerebras CS-3 — Phase 2 · quote only
- IBM z17 — Phase 2+ · quote only
- AMD Helios — Phase 2 · quote only
- d-Matrix — Phase 2–3 · quote only
Reserve Sovereign Compute Capacity
Secure priority access to AGICY's Phase 2 Cerebras CS-3 training clusters. Fine-tune and pre-train frontier models entirely within EU jurisdiction.
Secure Your SRA Allocation →References & Primary Sources
- Cerebras Systems — WSE-3 Chip Architecture and CS-3 Specifications
- Cerebras Inference Benchmarks — Llama 3.1 70B & 8B Throughput Results
- Groq LPU Technology — Architecture and Performance Claims
- SEC EDGAR — Cerebras Systems S-1 Filing and IPO Documentation (CBRS)
- AGICY Holdings Internal Business Plan, Sections 4.3 (Phase 2 Training) and 5.1 (Multi-Silicon Strategy), 2026.