Celeris-1-decision: The Smartest Decision Model Yet, 60% faster than Jev
Celeris-1-decision is a diffusion language model built for typed decisions. It takes text, structured data, or images, answers the yes/no, choice, or score questions you define with a probability for each outcome, and returns a one-sentence explanation on request. On jev-bench and typed decisions it beats Perplexity pplx-decider-v1.1-27b, Inception Mercury Decide, Jev 1.13.0 and OpenAI gpt-6-luna on both accuracy and Brier score, scoring 80.0% on jev-bench, at a 67 ms median response and $0.04 per million input tokens with output free. It returns a reason with each response, so the decisions it makes can be checked by a person, thresholded by code, and audited later.
Results
Celeris-1-decision has the highest accuracy and the lowest Brier score on jev-bench and typed decisions across all five decision APIs.
| Measure | celeris-1-decision | Perplexity pplx-decider-v1.1-27b | Inception Mercury Decide | Jev 1.13.0 | OpenAI gpt-6-luna |
|---|---|---|---|---|---|
| Full jev-bench accuracy (22 datasets, 3,210 items) | 0.800 | 0.790 | 0.753 | 0.731 | 0.690 |
| Full jev-bench Brier (lower is better) | 0.272 | 0.288 | 0.372 | 0.380 | 0.455 |
| Typed-decisions accuracy (400 × 5 questions) | 0.765 | 0.738 | 0.715 | 0.736 | 0.718 |
| Typed-decisions Brier | 0.079 | 0.173 | 0.270 | 0.149 | 0.204 |
On all 22 datasets of jev-bench (3,210 items) it scores 80.0% accuracy, against 79.0% for Perplexity pplx-decider-v1.1-27b, 75.3% for Inception Mercury Decide, 73.1% for Jev 1.13.0 and 69.0% for OpenAI gpt-6-luna. Its Brier score is 0.272, 28% lower than Jev (0.380) and 40% lower than gpt-6-luna (0.455), 6% lower than Perplexity (0.288) and 27% lower than Mercury Decide (0.372), and it wins 17 of the 22 datasets against Jev and gpt-6-luna.
Accuracy
Celeris-1-decision leads Perplexity by 1.0 point, Mercury Decide by 4.7 points and Jev 1.13.0 by 6.9 points on full jev-bench. On typed decisions it leads Perplexity by 2.7 points and Mercury Decide by 5.0. Against Jev and gpt-6-luna it wins 17 of the 22 jev-bench datasets.
Calibration
A decision API returns probabilities that callers threshold and act on, so their quality matters as much as the top answer. The Brier score measures it: the squared distance between the predicted distribution and the gold one. Lower Brier scores are better.
On typed decisions, where gold labels are soft distributions, celeris-1-decision's Brier of 0.079 is 47% below Jev's, 54% below Perplexity's, 61% below gpt-6-luna's and 71% below Mercury Decide's (0.270). On full jev-bench it is 6% below Perplexity, 27% below Mercury Decide, 28% below Jev and 40% below gpt-6-luna.
Latency
Median end-to-end latency, measured from a client in AWS us-east-1, was 67 ms for a one-question request and 81 ms for a fifty-question request. The other providers measured between 110 ms and 147 ms for one question and between 153 ms and 888 ms for fifty. Neither Celeris nor Jev serves inference from us-east-1, so every figure includes a cross-region network path. We chose us-east-1 because it is the most common region for cloud workloads, not because it favors any provider.
| Measure (AWS us-east-1) | celeris-1-decision | Perplexity | Mercury Decide | Jev 1.13.0 | gpt-6-luna |
|---|---|---|---|---|---|
| Median response, 1 question | 67 ms | 122 ms | 147 ms | 135 ms | 110 ms |
| Median response, 50 questions in one request | 81 ms | 888 ms | 460 ms | 153 ms | 175 ms |
Errors in the table are load rejections from the concurrency test, not malformed answers. Perplexity's are 429 responses from its 10 requests per second organisation limit.
Price
| Measure | celeris-1-decision | Perplexity | Mercury Decide | Jev 1.13.0 | gpt-6-luna |
|---|---|---|---|---|---|
| Price per 1M input tokens (output free) | $0.04 | $0.04 | $0.04 | $0.042 | $0.10 |
Price per token understates the comparison for decision workloads, where the unit that matters is a decision. A typical one-question request with a short text state is about 200 input tokens. At $0.04 per million tokens that is $0.000008 per decision, or about $8 per million decisions. The receipt example below, with an image at 560 soft tokens, is 818 input tokens, about $0.00003 per decision. Explanations do not change the price because output is free.
Deployment
Celeris-1-decision is a hosted API. There are no open weights and no fine-tuning in this release. The request format follows the systemone convention, so clients written for other decision endpoints can point at the Celeris base URL and add "x_celeris": {"explain": true} to receive explanations.
Methodology
All five providers received the same questions through their own decision endpoints and were scored by one scorer. gpt-6-luna's questions were translated to OpenAI's predicate, choice and score format. Perplexity ran at 2 requests in flight, under its 10 requests per second limit. Mercury Decide was called through Inception's System One-compatible endpoint.
Test sets. Every dataset is pinned to a fixed Hugging Face revision.
| Suite | Source | Items |
|---|---|---|
| Full jev-bench | Praveenrajus/jev-bench, all 22 configs | 3,210 |
| Typed decisions | LocalLLaMA/typed-decisions | 400 cases × 5 questions |
Metrics. Accuracy counts the top-probability option as the answer; a tie scores 1/k across the k tied options. Brier is the sum over options of (predicted − gold probability)², averaged over decisions; typed-decisions gold is a soft distribution, jev-bench gold a hard label. A failed request or a missing distribution scores as uniform over the options. Probabilities summing to within 5% of 1 are renormalised for every provider. Paired bootstrap 95% confidence intervals were computed per dataset.
Latency. A stdlib-only client on an AWS us-east-1 c7i.large measured warm single requests, cold connections, 1 to 50 questions per request, 200 to 32,000-character inputs, and 1, 4, 16 and 32 requests in flight. Figures are end-to-end from the client.
Questions people ask
What is celeris-1-decision? A hosted decision model from Celeris. It returns a probability for each typed question you ask about text, structured data, or an image, and a one-sentence explanation on request.
How accurate is it? 80.0% on the full jev-bench and 76.5% on typed decisions, the highest of the five hosted decision models we tested, with the lowest Brier score on both.
How fast is it? 67 ms median for one question, 81 ms for fifty, measured end to end from AWS us-east-1.
What does it cost? $0.04 per million input tokens, about $8 per million short decisions. Output tokens, including explanations, are free.
Does it work with my existing decision-model client? Yes. The request format follows the systemone convention. Point the client at the Celeris base URL; add the explain flag if you want reasons.
Example request
A receipt image and the claim it is attached to, with three yes/no questions and explanations on. The response below is verbatim from the live endpoint on Oct 8, 2026.
POST https://inference.celeris.ai/celeris-1-decision/v1/systemone
{
"model": "celeris-1-decision",
"state": { "claim": { "amount": 13.50, "category": "travel", "merchant": "cityride" } },
"questions": {
"matches": { "type": "noul", "instructions": "The receipt total matches the claimed amount." },
"travel": { "type": "noul", "instructions": "This is a travel expense." },
"altered": { "type": "noul", "instructions": "The receipt appears altered or edited." }
},
"x_celeris": {
"explain": true,
"images": { "receipt": "data:image/png;base64,..." },
"max_soft_tokens": 560
}
}
{
"model": "celeris-1-decision",
"answers": {
"matches": { "type": "noul", "noul": 0.94 },
"travel": { "type": "noul", "noul": 0.96 },
"altered": { "type": "noul", "noul": 0.04 }
},
"usage": { "input_tokens": 818, "output_tokens": 107 },
"x_celeris": {
"explanations": {
"matches": "The receipt total of $13.50 (base fare, booking fee, and tip) matches the claimed amount of 13.5 exactly.",
"travel": "The receipt is a 'Trip receipt' from Cityride for a 4.1 mi trip from Downtown to Market St, confirming it is a travel expense.",
"altered": "The receipt shows consistent fonts, spacing, and alignment throughout with no signs of digital editing or tampering."
}
}
}
818 input tokens, of which 560 are the image. Output tokens, including the explanations, are not billed.