Pay per token on the fastest model on Earth. Runs inside AWS and GCP, in your region - one API call, nothing to deploy.
Full-speed Celeris-1. Pay for what you use.
For teams running Celeris on the critical path at scale.
Prices in USD. Billed per million tokens; input and output metered separately.
Celeris-1 is the first model from Celeris: a general purpose language model that delivers near-GPT-5 level intelligence with up to 15x faster response times. It scores 75.9% on MMLU-Pro at a p50 latency of 158ms.
Celeris-1 is built from the ground up for low latency rather than adapted for it after the fact. We measure intelligence and speed on the same footing and publish the methodology, so you can verify the tradeoff yourself: 75.9% on MMLU-Pro at a p50 latency of 158ms.
Yes, and we publish everything. One harness, identical prompts and scoring across GPT-5, GPT-5-mini, Gemini 2.5 Flash, and Celeris-1. Every claim on this site links back to that data.
Voice, agents, and real-time applications. Anywhere waiting for a model breaks the experience, Celeris-1 is built to remove the wait.
Yes. Every plan runs the same full-speed Celeris-1: same weights, same latency, same quality. Speed is the product, not a premium tier, so what you test is exactly what you ship.
Per million tokens, input and output metered separately. No reserved capacity, no idle GPU charges, no surprises at the end of the month.
Enterprise plans can deploy in your VPC, in your region, so requests never leave your cloud's network. Talk to us about compliance and data residency requirements.
Celeris-1 is available starting today. Sign up, get an API key, and start building.