IonRouter
High-throughput, low-cost inference powered by IonAttention engine.
- Category
- AI & Machine Learning
- Launched
- March 12, 2026
- Stage
- Seed
- Pricing
- freemium
- Website
- ionrouter.io
IonRouter is a Y Combinator backed startup (W26) providing high throughput, low cost inference solutions for AI models. Its core technology is the IonAttention engine, a custom inference stack built from the ground up for NVIDIA Grace Hopper superchips. The engine multiplexes multiple models on a single GPU, swaps models in milliseconds, and adapts to traffic in real time. IonRouter claims throughput of 7,167 tokens per second on a single GH200 for Qwen2.5 7B, compared to 3,000 tok/s from a top inference provider. The platform supports any model—including finetunes, custom LoRAs, and open source models—with dedicated GPU streams, no cold starts, and per second billing. It offers an OpenAI compatible API, requiring only a one line change to switch from OpenAI. Pricing is per million tokens with no idle costs.
Discovered via
- Hacker News Launch posts · · Launch HN: IonRouter (YC W26) – High-throughput, low-cost inference