Cerebras
World's fastest AI inference with wafer-scale chips and 15x GPU speed
- Category
- AI & Machine Learning
- Launched
- July 23, 2025
- Stage
- Public
- Pricing
- freemium
- Website
- cerebras.ai

Cerebras Systems builds the world's fastest AI inference platform using its proprietary Wafer Scale Engine (WSE), a chip 58x larger than traditional GPUs. The company offers cloud, dedicated, and on premise deployment options for running large language models (LLMs) at unprecedented speeds. Cerebras recently launched Qwen3 235B, achieving 1,500 tokens per second with full 131k context support. The platform supports training, fine tuning, and serving on a single infrastructure, with drop in OpenAI API compatibility for easy integration.
Discovered via
- Hacker News Launch posts · · Cerebras launches Qwen3-235B, achieving 1.5k tokens per second
- Hacker News Launch posts · · AMD and Cerebras Launch AI Inference Solution