← Back to blog

Celeris-1-decision: The Smartest Decision Model Yet, 60% faster than Jev

Tom Hamer · October 8, 2026

Celeris-1-decision is a diffusion language model built for typed decisions. It takes text, structured data, or images, answers the yes/no, choice, or score questions you define with a probability for each outcome, and returns a one-sentence explanation on request. On jev-bench and typed decisions it beats Perplexity pplx-decider-v1.1-27b, Inception Mercury Decide, Jev 1.13.0 and OpenAI gpt-6-luna on both accuracy and Brier score, scoring 80.0% on jev-bench, at a 67 ms median response and $0.04 per million input tokens with output free. It returns a reason with each response, so the decisions it makes can be checked by a person, thresholded by code, and audited later.

Results

Celeris-1-decision has the highest accuracy and the lowest Brier score on jev-bench and typed decisions across all five decision APIs.

decision-benching harness 0cfd276 · one scorer for all providers · October 2026.
Measureceleris-1-decisionPerplexity pplx-decider-v1.1-27bInception Mercury DecideJev 1.13.0OpenAI gpt-6-luna
Full jev-bench accuracy (22 datasets, 3,210 items)0.8000.7900.7530.7310.690
Full jev-bench Brier (lower is better)0.2720.2880.3720.3800.455
Typed-decisions accuracy (400 × 5 questions)0.7650.7380.7150.7360.718
Typed-decisions Brier0.0790.1730.2700.1490.204

On all 22 datasets of jev-bench (3,210 items) it scores 80.0% accuracy, against 79.0% for Perplexity pplx-decider-v1.1-27b, 75.3% for Inception Mercury Decide, 73.1% for Jev 1.13.0 and 69.0% for OpenAI gpt-6-luna. Its Brier score is 0.272, 28% lower than Jev (0.380) and 40% lower than gpt-6-luna (0.455), 6% lower than Perplexity (0.288) and 27% lower than Mercury Decide (0.372), and it wins 17 of the 22 datasets against Jev and gpt-6-luna.

Accuracy

Celeris-1-decision leads Perplexity by 1.0 point, Mercury Decide by 4.7 points and Jev 1.13.0 by 6.9 points on full jev-bench. On typed decisions it leads Perplexity by 2.7 points and Mercury Decide by 5.0. Against Jev and gpt-6-luna it wins 17 of the 22 jev-bench datasets.

Calibration

A decision API returns probabilities that callers threshold and act on, so their quality matters as much as the top answer. The Brier score measures it: the squared distance between the predicted distribution and the gold one. Lower Brier scores are better.

On typed decisions, where gold labels are soft distributions, celeris-1-decision's Brier of 0.079 is 47% below Jev's, 54% below Perplexity's, 61% below gpt-6-luna's and 71% below Mercury Decide's (0.270). On full jev-bench it is 6% below Perplexity, 27% below Mercury Decide, 28% below Jev and 40% below gpt-6-luna.

Latency

Median end-to-end latency, measured from a client in AWS us-east-1, was 67 ms for a one-question request and 81 ms for a fifty-question request. The other providers measured between 110 ms and 147 ms for one question and between 153 ms and 888 ms for fifty. Neither Celeris nor Jev serves inference from us-east-1, so every figure includes a cross-region network path. We chose us-east-1 because it is the most common region for cloud workloads, not because it favors any provider.

Measure (AWS us-east-1)celeris-1-decisionPerplexityMercury DecideJev 1.13.0gpt-6-luna
Median response, 1 question67 ms122 ms147 ms135 ms110 ms
Median response, 50 questions in one request81 ms888 ms460 ms153 ms175 ms

Errors in the table are load rejections from the concurrency test, not malformed answers. Perplexity's are 429 responses from its 10 requests per second organisation limit.

Price

Measureceleris-1-decisionPerplexityMercury DecideJev 1.13.0gpt-6-luna
Price per 1M input tokens (output free)$0.04$0.04$0.04$0.042$0.10

Price per token understates the comparison for decision workloads, where the unit that matters is a decision. A typical one-question request with a short text state is about 200 input tokens. At $0.04 per million tokens that is $0.000008 per decision, or about $8 per million decisions. The receipt example below, with an image at 560 soft tokens, is 818 input tokens, about $0.00003 per decision. Explanations do not change the price because output is free.

Deployment

Celeris-1-decision is a hosted API. There are no open weights and no fine-tuning in this release. The request format follows the systemone convention, so clients written for other decision endpoints can point at the Celeris base URL and add "x_celeris": {"explain": true} to receive explanations.

Methodology

All five providers received the same questions through their own decision endpoints and were scored by one scorer. gpt-6-luna's questions were translated to OpenAI's predicate, choice and score format. Perplexity ran at 2 requests in flight, under its 10 requests per second limit. Mercury Decide was called through Inception's System One-compatible endpoint.

Test sets. Every dataset is pinned to a fixed Hugging Face revision.

SuiteSourceItems
Full jev-benchPraveenrajus/jev-bench, all 22 configs3,210
Typed decisionsLocalLLaMA/typed-decisions400 cases × 5 questions

Metrics. Accuracy counts the top-probability option as the answer; a tie scores 1/k across the k tied options. Brier is the sum over options of (predicted − gold probability)², averaged over decisions; typed-decisions gold is a soft distribution, jev-bench gold a hard label. A failed request or a missing distribution scores as uniform over the options. Probabilities summing to within 5% of 1 are renormalised for every provider. Paired bootstrap 95% confidence intervals were computed per dataset.

Latency. A stdlib-only client on an AWS us-east-1 c7i.large measured warm single requests, cold connections, 1 to 50 questions per request, 200 to 32,000-character inputs, and 1, 4, 16 and 32 requests in flight. Figures are end-to-end from the client.

Questions people ask

What is celeris-1-decision? A hosted decision model from Celeris. It returns a probability for each typed question you ask about text, structured data, or an image, and a one-sentence explanation on request.

How accurate is it? 80.0% on the full jev-bench and 76.5% on typed decisions, the highest of the five hosted decision models we tested, with the lowest Brier score on both.

How fast is it? 67 ms median for one question, 81 ms for fifty, measured end to end from AWS us-east-1.

What does it cost? $0.04 per million input tokens, about $8 per million short decisions. Output tokens, including explanations, are free.

Does it work with my existing decision-model client? Yes. The request format follows the systemone convention. Point the client at the Celeris base URL; add the explain flag if you want reasons.

Example request

A receipt image and the claim it is attached to, with three yes/no questions and explanations on. The response below is verbatim from the live endpoint on Oct 8, 2026.

POST https://inference.celeris.ai/celeris-1-decision/v1/systemone

{
    "model": "celeris-1-decision",
    "state": { "claim": { "amount": 13.50, "category": "travel", "merchant": "cityride" } },
    "questions": {
        "matches": { "type": "noul", "instructions": "The receipt total matches the claimed amount." },
        "travel":   { "type": "noul", "instructions": "This is a travel expense." },
        "altered": { "type": "noul", "instructions": "The receipt appears altered or edited." }
    },
    "x_celeris": {
        "explain": true,
        "images": { "receipt": "data:image/png;base64,..." },
        "max_soft_tokens": 560
    }
}
{
    "model": "celeris-1-decision",
    "answers": {
        "matches": { "type": "noul", "noul": 0.94 },
        "travel":       { "type": "noul", "noul": 0.96 },
        "altered": { "type": "noul", "noul": 0.04 }
    },
    "usage": { "input_tokens": 818, "output_tokens": 107 },
    "x_celeris": {
        "explanations": {
            "matches": "The receipt total of $13.50 (base fare, booking fee, and tip) matches the claimed amount of 13.5 exactly.",
            "travel":    "The receipt is a 'Trip receipt' from Cityride for a 4.1 mi trip from Downtown to Market St, confirming it is a travel expense.",
            "altered": "The receipt shows consistent fonts, spacing, and alignment throughout with no signs of digital editing or tampering."
        }
    }
}

818 input tokens, of which 560 are the image. Output tokens, including the explanations, are not billed.