Hardware · Cerebras · Pre-COD
Wafer-scale AI supercomputer
Phase 2 evaluation at AGICY Vasilikos — silicon optionality, not endorsement
CS-3 is one 300 mm wafer as a single engine: 4 trillion transistors, 900,000 AI cores, 44 GB on-chip SRAM. On 18 August 2026 Cerebras announced CS-4— a Nexus rack with three WSE-3 Turbo wafers, not a new WSE-4 die. Phase 1 capital planning remains Galaxy. Neither CS-3 nor CS-4 is an owned AGICY fleet.
- WSE-3 / WSE-3 Turbo · 300 mm wafer
- 44 GB on-chip SRAM per wafer
- CS-4 · 3 wafers · Nexus rack
- Phase 2 · announced · not ordered
- No Cerebras endorsement
- Vendor-claimed figures under diligence
- Not Phase 1 capital plan
- Pre-COD · roadmap optionality
Elevation · CS-3 rack · vendor photo · not CS-4CS-4 — Nexus rack, not WSE-4
Cerebras announced CS-4 on 18 August 2026: the first Nexus rack-scale system, built from 3 WSE-3 Turbo (WSE-3T) processors. AGICY has no CS-4 order, no public price, and no endorsement.VERIFIED
Each WSE-3 Turbo wafer still lists 4T transistors, 900,000 cores, and 44 GB on-wafer SRAM — the same counts as WSE-3. Cerebras bills WSE-3T as roughly 2× AI compute and memory bandwidth per wafer(250 vs 125 PFLOPS; 43.2 vs 21.6 PB/s). The IR comparison table below is one CS-3 wafer versus a three-wafer CS-4 rack, so 750 vs 125 PFLOPS is not a like-for-like chip jump.VENDOR-CLAIMED
| Metric (Cerebras IR, 18 Aug 2026) | CS-3 (one wafer) | CS-4 (three wafers) |
|---|---|---|
| AI compute | 125 PFLOPS | 750 PFLOPS |
| Memory bandwidth | 21.6 PB/s | 129.6 PB/s |
| On-chip fabric bandwidth | 26.7 PB/s | 160.5 PB/s |
| System I/O bandwidth | 1.2 Tbit/s | 7.2 Tbit/s |
| I/O latency | 5 µs | 2 µs |

Cerebras' own axes: interactivity is tokens per second per user; throughput / token capacity is total tokens per second per megawatt. The callouts are up to 2× (X-axis) and up to 10× (Y-axis) versus CS-3. GPU is shown as high total TPS/MW at low TPS/user. This is vendor marketing, not an AGICY productivity or CapEx input.VENDOR-CLAIMED
- Up to 2× interactivity (TPS/user) vs CS-3 (Cerebras); Up to 10× token capacity (total TPS per MW) vs CS-3 (Cerebras)VENDOR-CLAIMED
- Up to 30× vs GPU systems on Cerebras’ chosen comparisons. More than 4,400 TPS/user on GPT-OSS-120B (Cerebras; varies by config)VENDOR-CLAIMED
- First CS-4 shipments begin this quarter (Cerebras, 18 Aug 2026). Independent coverage (TNW, The Next Platform, The Register) reads WSE-3T as a clock/power bump on the WSE-3 die, not a new wafer generation — press reading, not an AGICY lab result.DESIGN TARGET
- Cerebras also describes disaggregated inference (prefill on Helios / Trainium, decode on CS-4). That is vendor architecture talk for the Phase 2 eval queue — not a partnership or campus design.VENDOR-CLAIMED
Vendor sources: Cerebras IR release · CS-4 blog · CS-4 datasheet.
Next: CS-3 video briefing →
CS-3 video briefing
Four-chapter briefing on WSE-3 architecture, on-chip SRAM, training modalities, and multi-silicon complementarity with Tenstorrent Galaxy. Footage and narration are CS-3 — CS-4 is covered in the announcement sheet above.
Wafer-scale AI supercomputers
Cerebras solves a different problem than open RISC-V modular accelerators — one giant SRAM-rich die for frontier training and ultra-high per-user throughput.
Most AI silicon is a network of chips — GPUs, LPUs, or RISC-V accelerators linked by fabric. Cerebras takes the opposite bet: keep the active model on a single wafer-scale engine so data never crosses chip boundaries during compute. That is a supercomputer architecture thesis, not a cheaper Galaxy substitute.VERIFIED
Tenstorrent's open RISC-V ecosystem targets modular, air-cooled inference economics and ISA sovereignty. Cerebras targets memory-wall elimination and extreme tokens-per-user on large models. AGICY keeps both on the map because the workloads diverge — not because one vendor "wins." VERIFIED
Next: WSE-3 architecture →
WSE-3: wafer-scale architecture
Third-generation Wafer Scale Engine — an entire 300 mm wafer as one chip.
The WSE-3 uses a full 300 mm wafer as a single die — 4 trillion transistors on TSMC 5 nm, 900,000 sparse-linear-algebra cores, 46,225 mm² wafer area (not a Galaxy/Blackhole figure). During active compute, tensors stay on-wafer rather than syncing across a multi-GPU cluster.VENDOR-CLAIMED
- 16RU liquid-cooled form factor per CS-3 system (vendor system page)VENDOR-CLAIMED
- ~23 kW TDP per unit (facility PUE 1.20 planning target)DESIGN TARGET
- SwarmX fabric for multi-wafer scale-outVENDOR-CLAIMED


On-chip SRAM & memory bandwidth
Forty-four gigabytes of wafer SRAM with 21 PB/s internal bandwidth — the memory-wall counter-argument.
44 GB SRAM on wafer with 21 PB/s internal bandwidth — no external HBM required for models that fit on chip. That removes the memory wall that limits GPU inference for large transformer layers.VENDOR-CLAIMED
Deep dive: Cerebras CS-3: Sovereign Training.
CS-3 vs Galaxy — different problems
Tenstorrent and Cerebras are both NVIDIA alternatives. They do not solve the same problem.VERIFIED
Cerebras CS-3
- Wafer-scale engine: one giant SRAM-rich die for frontier training and ultra-high per-user throughputVENDOR-CLAIMED
- ~$1.5M/node class, liquid-cooled 16RU — denser facility fitVENDOR-CLAIMED
- Phase 2 evaluation / roadmap — complementary, not a Galaxy substitute
Tenstorrent Galaxy
- Modular network-of-chips: 32× Blackhole per 6U server, open RISC-V software path, GDDR6, Ethernet scale-outVENDOR-CLAIMED
- ~$110K/server class, 9 kW air-cooled — Phase 1 capital and P&L baseline
- Sovereign batch and civil inference at fleet scale (pre-COD)
AGICY routes by workload: Galaxy for sovereign batch and civil inference at fleet scale; CS-3 where wafer-scale training or latency gates win. Neither vendor endorses AGICY's model — both are silicon supply options under diligence.VERIFIED
Tenstorrent Galaxy hardware page → · IBM z17 → · AMD Helios → · d-Matrix Corsair →
Inference throughput
Vendor-claimed tokens-per-second on 70–120B-class models — independent validation pending before any CapEx commit.
| Model class | CS-3 (vendor-claimed) | Groq LPU (ref.) |
|---|---|---|
| Llama 3.1 70B | 1,500–2,100 tok/s | ~250 tok/s |
| gpt-oss-120B | 3,000+ tok/s | — |
| Time to first token (p50) | ~12 ms | Varies |
Training scale & modalities
Weight Streaming decouples compute from model storage — frontier training configurations across clustered CS-3 systems.
Weight Streaming enables training configurations up to 24 trillion parameters across clustered CS-3 systems (vendor-claimed). cerebras.pytorch compiles standard PyTorch for text, vision, diffusion, and multimodal workloads.VENDOR-CLAIMED
Vasilikos Phase 2 facility fit
Liquid cooling and 16RU racks — planned in Phase 2/3 expansion halls, not Phase 1 air-cooled Galaxy floors.
CS-3 requires liquid coolingand 16RU racks — planned in Phase 2/3 expansion halls with hybrid air + liquid plant. Mediterranean free-cooling and the 42 MW→200 MW power roadmap support high-density wafer-scale pods under EU jurisdiction, subject to technology committee and Financial Close.DESIGN TARGET

Sovereign AI Campus briefing →
Next: Diligence & honesty →
Diligence & honesty
What AGICY claims — and what it does not — about Cerebras on campus.
- Vendor-claimed throughput pending independent SC-class benchmarksDESIGN TARGET
- Liquid-cooling retrofit and technology-committee gates before CapExDESIGN TARGET
- Supplier, software maturity, and facility density risks — standard pre-production diligenceDESIGN TARGET
AGICY relevance
Multi-silicon sovereign posture (pre-COD) — wafer-scale optionality beside the Galaxy Phase 1 anchor.
Related research
Interested in Leasing?
Lease or finance CS-3 or CS-4 units on AGICY premises, or at your site if it meets our T&Cs. Free estimate for integration and delivery (quote-only — Cerebras has not published a CS-4 list price). Pre-COD; Leasing ≠ SRA / CY tokens.