AI DevelopmentFramework6 min readPublished September 28, 2026

4 publishers · 7 routes · output tokens free · a classifier that reads like a language model

Decision Models: AI That Returns Probabilities, Not Text

Four publishers now list decision models on OpenRouter that return probabilities for typed questions, not prose, at free to $0.05 per million input tokens.

DA
Digital Applied Team
Research and practical guidance
Catalog checkedSeptember 28, 2026
Routes7 decisions routes

A small new category of AI model answers questions without writing anything. You send it some text, such as a support message or an agent’s transcript, plus a list of typed questions: is this urgent, which team should get it, did the agent loop. It returns a probability for each answer and nothing else. On September 28, 2026, OpenRouter listed seven of these routes from four publishers. Three of the four publishers arrived in the previous four days.

For anyone running agents, the appeal is cost and speed. Output tokens are free, because there is no output text, and each answer takes a single pass through the model. The question is whether the answers are good enough to replace the language model many teams now use as a judge. The vendors say yes; the evidence so far is their own.

Key takeaways
  1. 01
    Decision models return probabilities for typed questions and generate no text.Three answer types: yes/no, a choice from options you define, or a position on a scale. Your code decides what to do with them.
  2. 02
    Seven routes from four publishers, priced from free to $0.05 per million input tokens.Output is free on every route. Upstage lists Solar Decide at $0.10 on its own console; the OpenRouter route shows it at half that.
  3. 03
    Every performance claim so far is vendor-run.Respan ranks Span-01 first of nine models on its headline behaviour benchmark, but its per-domain table puts GPT-6 Sol ahead. Upstage published no benchmark.
  4. 04
    They run on a separate API, so chat SDKs will not call them.OpenRouter serves them through an alpha Decisions API with its own request and response shapes.

01 — The ideaWhat a decision model is

A normal language model writes an answer word by word, and if you want a label you parse it out of the text. A decision model skips the writing. According to OpenRouter’s route page, a request carries a state, which can be a string, an object or an array, and a set of questions. Each question has a type: noul for a yes/no probability, choice for a pick from options you define, or score for a position on an ordered rubric. The model returns the probabilities and your code acts on them.

These models do not run on the familiar chat endpoint. OpenRouter serves them through a separate, alpha Decisions API, and its page warns that chat-completions SDKs will not work with it. The format started with TypeSafe’s Jev, which we covered at its launch. Upstage describes Solar Decide as a System One endpoint, the name TypeSafe gave that format.

noul
Yes or no
One probability

Is this message urgent? Does this reply contain a secret? The model returns the probability of yes.

Guardrails, triage
choice
Pick one
A probability per option

Which team should handle this ticket? Which tool should the agent call next? Options are defined in the request.

Routing
score
Place on a scale
A position on a rubric

How complete is this answer, on a rubric you write? Useful for grading and ranking.

Evaluation

02 — The censusThe routes on sale on September 28

We filtered OpenRouter’s model catalog, read at 21:28 UTC on September 28, for routes whose output type is decisions. The listing dates are OpenRouter’s; the release dates are from each vendor’s own page where one exists.

Source: OpenRouter model catalog, September 28, 2026, 21:28 UTC; Upstage console and Respan announcement for vendor details. Input price per million tokens on the route; output is free on all.
Model and routeInputWhat it is
TypeSafe Jev 1.13 (typesafe/jev-1.13)$0.042The first on the endpoint: announced by TypeSafe September 15, listed on OpenRouter September 18. 32,000-token context. Returns yes/no, choice and score answers.
Upstage Solar Decide (upstage/solar-decide)$0.05Released September 22 in beta on Upstage's console at $0.10; the route shows 50% off. Built on Solar Mini 4, 512K context, same request format as Jev.
Respan Span-01 (respan/span-01)$0.02Announced September 24. Scores behaviours you define in a conversation as present, absent or not observable.
Respan Span-01 Lite (two routes)FreeA lighter tier of Span-01 for high-volume monitoring; one route is marked :free.
Kev-4B (jaredpalmer/kev-4b)$0.042Jared Palmer's open-weight 4B model, a LoRA adapter on Qwen3.5-4B-Base with weights on Hugging Face. 8,192-token context.

The seventh route is ~typesafe/jev-latest, an alias that points at the current Jev version. As with any alias, pin a version for anything you measure. Solar Decide is the notable newcomer: it comes from an established model maker rather than a startup, and Upstage lists it as beta, with a 512K-token context that can take a whole document as the state.

03 — The claimsWhat the vendors claim, and who measured it

Upstage describes Solar Decide’s probabilities as calibrated and says each decision is one forward pass. It publishes no benchmark, so there is nothing to compare yet.

Respan, which sells tools for tracing and evaluating AI agents, published two benchmarks with its September 24 announcement. Span-01 is built for one job: you write a plain description of a behaviour, such as user frustration or an agent misusing a tool, and it returns the probability that the behaviour is present, absent or not observable in a conversation. On the headline benchmark Respan ranks Span-01 first of nine models at 0.843 overall F1, ahead of GPT-5.6 Terra (0.837) and GPT-6 Luna (0.815); GPT-6 Sol is not in that set. Respan’s results by behaviour domain, which do include Sol, all run by Respan:

Source: Respan, Span-01 announcement, September 24, 2026. Vendor-run; four of seven behaviour domains shown, plus the overall score.
Behaviour domainSpan-01JevGPT-6 Sol
Jailbreak and prompt injection0.7790.7520.878
Hallucination and grounding0.7960.6730.903
Agent and tool reliability0.8450.6910.861
Response quality0.7710.6910.956
Overall0.8060.7160.885

Two things stand out. Span-01 beats Jev, and Sonnet 5 (0.719 overall), in every domain shown. And GPT-6 Sol, a full frontier model, scores higher than all of them in six of Respan’s seven domains, by the widest margin on response quality; Span-01 leads only on privacy and secrets. Respan concedes the point in its own write-up. The pitch is not that a decision model is as good as the best judge, but that it is close enough at a fraction of the price.

Respan also ran 11 decision models through a test in which Span-01 itself did the judging, measuring accuracy, consistency, resistance to prompt injection and calibration. Jev 1.13 ranked first at 0.932 accuracy and Kev-4B scored 0.833. Treat that ranking with care: the judge was the vendor’s own model.

Not observable is not absent

Span-01’s third answer matters. Respan’s documentation gives the example of a conversation that ends at the assistant’s reply: it cannot show whether the customer accepted the fix, so the right answer is “unknown,” not “absent.” A monitoring rule that treats a high p_not_observable as a pass will miss real problems.

04 — The costThe cost case against an LLM judge

An illustrative workload shows why teams are interested. Suppose an agent platform checks one million conversations a month, each about 2,000 tokens long. That is two billion input tokens. The volume is invented for the example; the prices are the listed route prices.

  • Span-01 at $0.02 per million input tokens: $40 a month, with no output charge.
  • Solar Decide at the $0.05 route price: $100 a month.
  • A Claude Sonnet 5.5 judge at $2 per million input tokens: $4,000 for input alone, before any output or thinking tokens.

The saving is real only if the cheaper model’s mistakes cost less than the difference. For many checks, such as routing a ticket or flagging a likely prompt injection for review, a small error rate is fine. For a decision that blocks a customer or triggers a refund, it may not be. Our post on model routers and classifier overhead covers the same trade-off for routing.

05 — The testHow to test a decision model on your own work

  1. Label a few hundred real cases. Use conversations or tickets where you know the right answer. Vendor benchmarks do not use your definitions.
  2. Set thresholds from that data. A probability is only useful with a cut-off, and the right cut-off depends on whether a false alarm or a miss costs more.
  3. Send the uncertain middle to a stronger check. Respan’s own guidance is to pass borderline cases to a person or a frontier model. That keeps most of the saving and limits the damage from errors.
  4. Rephrase your questions and re-run. A good decision model should give the same answer to the same question asked two ways. Respan measures this as a flip rate; measure it yourself.

If your agents already produce structured outputs, the structured output reliability guide covers the adjacent problem of getting typed answers out of a normal model.

06 — ConclusionDecision models suit narrow, high-volume checks

You monitor agent conversations at scale
Trial Span-01 or Span-01 Lite against your labelled cases, with borderline results sent to a frontier model.
Behaviour monitoring
You route or triage tickets and requests
Compare Jev and Solar Decide on choice questions. Solar Decide's long context suits whole documents.
Routing
The decision blocks a user or moves money
Keep a frontier model or a person in the loop. On Respan's per-domain benchmark, GPT-6 Sol still outscores Span-01 and Jev overall.
High-stakes checks
What to do this week

Pick one high-volume check your agents already make, label a few hundred cases, and compare a decision model against your current judge

The category is less than two weeks old on OpenRouter, and the only benchmarks come from a vendor. The price gap is large enough to justify a test anyway. Start with a check where an occasional error is cheap, measure it on your own data, and move to higher-stakes checks only once the error rate is known.

Digital Applied

Check every agent turn without paying frontier prices.

We design the evaluation and guardrail layers for AI agents, from labelled test sets to the thresholds that decide when a person steps in.

Agent evaluationGuardrail designCost modelling
Your next project

Agent checks you can trust

  • →Labelled cases from your own traffic
  • →Thresholds set from data
  • →A fallback for the uncertain middle
Questions and answers

The questions we get about decision models

A model that answers typed questions about a piece of text, such as yes or no, a choice between options, or a score, by returning probabilities instead of writing text. Each answer takes one pass through the model, and output tokens are free.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading

AI Development

Fireworks Ember-1: Kimi K3 Quality With Fewer Tokens?

Fireworks tuned Kimi K3 into Ember-1 and says it matches K3 with 35 to 50% shorter reasoning at the same price. The rows it loses, and the preview caveat.

September 23, 2026 · 5 minRead
AI Development

Claude Sonnet 5.5: Pricing, Benchmarks and Safeguards

Claude Sonnet 5.5 keeps Sonnet 5's $2/$10 price, scores 70.6% on Terminal-Bench 4.0 and adds cyber fallbacks. What the 30% saving is measured on.

September 28, 2026 · 9 minRead
AI Development

Gemini 3.8 Flash TTS: Voice Cloning and a Price That Doubles

Gemini 3.8 Flash TTS is generally available with voice cloning from a 30-second sample. The promotional price ends December 31 and doubles on January 1, 2027.

September 23, 2026 · 5 minRead
AI Development

GPT-6 Sol vs Claude Opus 5.5: Cost per Task and Benchmarks

Artificial Analysis ran GPT-6 Sol and Claude Opus 5.5 on the same tests. Sol is cheaper up to a point; Opus 5.5 at its default outscores Sol at max.

September 22, 2026 · 6 minRead
AI Development

What People Let AI Agents Access: 2026 Survey Numbers

A July 2026 survey of 5,067 US adults: 41% of AI users have tried an agent and 32% have let AI act without a final sign-off. Every figure with its base.

September 16, 2026 · 6 minRead
AI Development

After AI Context Compaction, Which Instructions Survive?

Check whether an AI agent follows the right instructions after context compaction. Use behavioral probes for task scope, permissions, evidence and progress.

September 12, 2026 · 6 minRead