Cerebras CS-3 Wafer-Scale Engine: 4 trillion transistors, 44GB SRAM, Llama 4 Maverick at 1500 tokens/sec
In one sentence Cerebras unveils the CS-3, its third-generation wafer-scale engine featuring 4 trillion transistors and 44 GB of on-chip SRAM, running Llama 4 Maverick at 1500 tokens per second on a single chip. First commercial deployment is live in the UAE AI cloud.
Picture a regular computer chip as a single apartment in a large building. Most AI chips are like that — small, numerous, and constantly passing messages to each other through hallways and elevators. Cerebras did something radical: instead of cutting a silicon wafer into many small chips, they use the entire wafer as one giant chip.
The result is the CS-3, a processor the size of a dinner plate, packed with 4 trillion transistors and 44 gigabytes of ultra-fast memory built directly into the chip itself. No waiting for data to travel from external RAM — everything the model needs is already on board.
Why does this matter? When an AI model like Llama 4 Maverick answers a question, it has to read billions of numbers (its weights) from memory on every single step. With a normal GPU setup, this is slow because the memory is far away. With the CS-3, all that memory is right there, on the chip, and it answers at 1500 tokens per second — roughly 5 to 10 times faster than a typical high-end GPU cluster.
The first real-world customer is the UAE AI national cloud, signaling that this technology is aimed at governments and large enterprises that need ultra-fast AI responses at scale — think real-time analytics, financial modeling, and large-scale language processing.
It is not a consumer product. But it challenges the idea that NVIDIA GPUs are the only serious option for AI inference, and that competition is good for the entire industry.
Companies
Cerebras
Tools
CS-3 Cerebras Inference
Tags
Sources