NEW Introducing Celeris-1 Magnus

Speed changes what's possible.

Building the world's fastest language models, powered by diffusion.

VOICE AGENTLIVE TURNS
Caller

Can I move my flight to Thursday morning?

Agent · turn 158 ms

There's a 9:40 am at the same fare. Want me to book it?

Caller

Yes, do it.

Agent · turn 142 ms

Booked. Confirmation is in your inbox.

Celeris-1 · streamingREAL MODEL SPEED
throughput 0 tok/selapsed 0 ms
AGENT LOOP50 CALLS
plan refactor 61 ms
grep "rateLimit" src/ 48 ms
read routes.ts · 42 lines 57 ms
edit src/api/routes.ts 64 ms
run tests · 14 passed 52 ms
task complete 3.4 s total

The fastest model ever measured.

Complete answers in 158 milliseconds, at 75.9% MMLU-Pro accuracy - up to 81.4% with reasoning enabled. The next fastest frontier model takes over a second.

Explore Celeris-1 →
Output speed · higher is betterTokens / sec
Celeris-1
5.6×
2,038
Gemini 3.7 Flash (high)
362
GPT-5.6 Luna (max)
141
DeepSeek V4 Pro (max)
72
Claude Fable 5
71
Grok 4.6 (high)
62

Figure 4. Output tokens per second across frontier models. Read the research →

Beyond the current architecture.

ARCHITECTURE

Language models built to generate and refine information in parallel.

GENERATION

Diffusion starts with a rough representation of the response and progressively refines it into language.

INFERENCE

Parallel refinement enables intelligence to move at unprecedented speed.

Explore our research →

When intelligence becomes real-time, everything changes.

Voice

Conversations without waiting for the model to respond.

Agents

Systems that reason, act, and adapt continuously.

Interaction

AI that responds at the pace of human input.

Autonomy

Systems that perceive, reason, and act in real time.

Intelligence at the speed of thought.

Start building with Celeris.

Experience the new paradigm of inference.

Get Started See the benchmarks