Lamb Labs
Model-first AI inference: diffusion post-training now, custom silicon later, targeting 20,000+ tokens per second.
- Category
- AI & Machine Learning
- Launched
- August 3, 2026
- Stage
- Y Combinator (S26)
- Pricing
- paid
- Founder
- Niki Kotecha, Thomas Lanning
- Website
- lamb-labs.com



The AI chip industry has a familiar reflex: when inference is too slow, add more cores, more memory bandwidth, more silicon. Lamb Labs, a Y Combinator Summer 2026 company, is trying the opposite. Their thesis is that the bottleneck isn't the chip at all — it's the model's sequential nature. Autoregressive models emit one token at a time, forcing the hardware to wait on memory. Lamb Labs' answer is to convert those models into diffusion architectures that decode in parallel, then build silicon to match. The claim: 20,000+ tokens per second and 63x higher intelligence per watt than a GPU like the RTX 6000 Ada. That's a bold target, but the more interesting story is the path — and the wedge that might make it work.
Discovered via
- Y Combinator Launches · · Lamb Labs