How we measure it, how we build it, and what a low-latency model unlocks.
We ran the complete GSM8K, IFEval and DROP benchmarks - 11,395 questions per model, reasoning off everywhere - against gpt-5, gpt-5-mini, the Gemini Flash models and mercury-2. celeris-1 averages 85.3%, within 1.9 points of gpt-5, while answering in about 0.1 seconds.
Our objective is to maximize useful intelligence delivered per unit of time. Why we're rethinking training and architecture from the ground up - hybrid autoregressive-diffusion decoding and automated architecture search - instead of only optimizing the status quo.
Measured from a client colocated with each provider, Celeris delivers 1,596 tokens per second end-to-end - 8% ahead of Cerebras and 3.5× ahead of Groq serving gpt-oss-120b - and why end-to-end wall clock is the one basis every provider can be compared on.
On an Artificial-Analysis-style workload, celeris-1 generated a median of 1,664 tokens per second - 5.1× the next-fastest API we measured - and returned complete responses to 1,000-token prompts in a median of 0.58 seconds. The full results, the harness, and the limitations.
We ran Celeris against GPT-5, GPT-5 mini, the Gemini Flash models, and Mercury 2 on the full MMLU-Pro test set - one identical harness, quality and latency from the same run. The methodology, the numbers, and how we separated model time from the network.