Building the world's fastest language models, powered by diffusion.
Can I move my flight to Thursday morning?
There's a 9:40 am at the same fare. Want me to book it?
Yes, do it.
Booked. Confirmation is in your inbox.
Complete answers in 158 milliseconds, at 75.9% MMLU-Pro accuracy - up to 81.4% with reasoning enabled. The next fastest frontier model takes over a second.
Explore Celeris-1 →Figure 4. Output tokens per second across frontier models. Read the research →
Language models built to generate and refine information in parallel.
Diffusion starts with a rough representation of the response and progressively refines it into language.
Parallel refinement enables intelligence to move at unprecedented speed.
Conversations without waiting for the model to respond.
Systems that reason, act, and adapt continuously.
AI that responds at the pace of human input.
Systems that perceive, reason, and act in real time.
Within two points of GPT-5 at 15x lower latency, on the complete benchmark suites.
2026 ComparisonFor developers building voice, agents, search, and other ultra-low latency AI systems.
2026 BenchmarkTransparent benchmarking methods and real-world performance testing.
2026Experience the new paradigm of inference.