Groq launches GroqCloud 2.0: LPU Gen3, 2000 tokens/sec, and Frankfurt European data center
In one sentence Groq releases GroqCloud 2.0 with third-generation LPU chips delivering 2000 tokens/sec on Llama 4.1 Maverick, adds Function Calling GA, JSON mode, streaming tool use, batch pricing, and opens a European data center in Frankfurt.
Think of inference speed like a typist: most AI services type at a reasonable pace, but Groq has built special hardware that types entire essays in the time others finish a sentence. That is the core promise of Groq's LPU chip — a processor designed specifically for running AI language models, not adapted from graphics cards like most competitors use.
With GroqCloud 2.0 and the third-generation LPU, Groq pushes that speed to 2000 tokens per second on Meta's Llama 4.1 Maverick model. In practical terms, that means full paragraphs of text generated nearly instantly, which matters a lot for real-time applications like chatbots, coding assistants, or voice interfaces.
Beyond raw speed, this release adds several features that developers have been waiting for: Function Calling is now stable and ready for production use, meaning AI models can reliably trigger actions in external systems. JSON mode ensures the model always returns properly structured data. Streaming tool use lets complex AI agents work more responsively without waiting for a full response before acting.
For cost-sensitive workloads, a new batch inference pricing tier offers significant discounts for jobs that do not need immediate results.
Perhaps the most strategically important addition is a data center in Frankfurt, Germany. European companies dealing with privacy regulations now have a compliant option to use Groq's fast inference without sending data outside the EU.
Overall, GroqCloud 2.0 strengthens Groq's position as the go-to platform for developers who want the fastest and most affordable way to run open-source models in production.
Companies
Groq
Tools
GroqCloud, LPU
Tags
Sources