Cost Optimization
The four-cent decision: stop paying essay prices for yes-or-no AI
A new kind of AI model answers with a typed decision, not prose, for about four cents per million tokens. How much of your AI work fits, and when it is worth a new vendor.
Satori Canton
October 11, 2026 · 12 min read
Every company running AI in production has a ticket router somewhere. A message arrives, a model reads it, and the model decides where it goes: billing, technical, sales. Most of those routers were built the same way. Someone wrote a prompt for a capable general model, asked it to answer with one word, and wrote code to pull that word out of whatever came back.
That works, and it is a strange way to buy a decision. A general model is built to write. When you ask it which of three queues a ticket belongs in, it is generating text that happens to contain an answer, and you pay for the generating.
In September a startup called TypeSafe released a model that does not write at all. Jev takes a piece of text and a set of typed questions, and returns typed answers: which option, with a probability for each; a score on a scale you define; the probability that a statement is true. It costs $0.042 per million input tokens, and output is free. Four cents reads a million tokens, enough for about 1,400 routine decisions of the size we price below.
Three weeks later OpenAI opened a public beta of a Decisions API that works the same way. A day after that, Anthropic released Claude Haiku 5.5 and described it as built for classification, routing and extraction. On October 9, TypeSafe announced an $870 million Series A at a $7.5 billion valuation, led by Andreessen Horowitz.
Our October 9 issue covered OpenAI and Anthropic and left out the model that started the category. This article corrects that, and asks the question a budget owner needs answered: how much of your AI work is actually a decision, and what should it cost?
What a decision model is
A decision model answers three kinds of question, and only those three. The names differ by vendor; the shapes do not.
| Question type | What it returns | Jev | OpenAI Decisions |
|---|---|---|---|
| Pick one from a list | The choice, a probability for each option, a confidence | Choice | choice |
| Rate on a scale | A score, a probability for each level, a confidence | Score | score |
| Is this true? | The probability that it is | Noul | predicate |
You can ask many questions about the same text in one call, and each is answered independently. In TypeSafe's own example, a single support message is asked which team should handle it, how frustrated the customer is, and whether it is urgent, and the answer comes back as data: technical at 85%, "frustrated but civil" at 100%, urgent. There is nothing to parse, so there is nothing to parse wrong. The answer can only be one of the values you allowed.
That is the first advantage, and engineers feel it before finance does. The second is speed: TypeSafe says most calls finish in about 100 milliseconds, and OpenAI says its version decides about 10 times faster than the same model through its regular API. The third is the probability. A general model gives you an answer. A decision model gives you an answer and how sure it is, and your code can treat "billing, 97%" differently from "billing, 52%".
What a decision model will not do is write. It does not draft the reply, summarize the thread, explain its reasoning or extract a value it has not been offered as an option. TypeSafe's documentation is plain about this: Jev "does not generate text, write code, or hold a conversation." A decision model sits beside your general models, not in place of them.
The price of a decision
Here is what one million routine decisions cost at list prices on October 10, for a call that sends 700 tokens (a ticket plus three questions) and, on a general model, gets 50 tokens back.
| Model | Price per million tokens, in / out | One million decisions |
|---|---|---|
| Claude Opus 5.5 | $4 / $20 | $3,800 |
| Claude Sonnet 5.5 or OpenAI GPT-6.1 Sol | $2 / $10 | $1,900 |
| Gemini 3.5 Flash-Lite | $0.30 / $2.50 | $335 |
| Claude Haiku 5.5 or GPT-6 Luna | $0.10 / $0.50 | $95 |
| OpenAI Decisions API (GPT-6 Luna) | $0.10 / no output charge | $70 |
| TypeSafe Jev | $0.042 / no output charge | $29 |
Sources: Anthropic, OpenAI, Google, TypeSafe. Tokenizers differ by vendor, and a general model that reasons before answering bills its thinking as output, so treat the general-model rows as floors.
Two things stand out. The first is the size of the drop from where most routers were built. Moving a decision from a mid-tier model to any of the bottom three rows cuts its cost by 95% or more. The second is how close together the bottom three rows are. Jev is less than half the price of OpenAI's Decisions API, but the difference is $41 per million decisions. That matters at a billion decisions. At ten million a month it is a rounding error next to the first saving.
TypeSafe named Jev after William Stanley Jevons, the economist who observed that using coal more efficiently led to more of it being burned, not less. Expect the same here. At $29 per million, a decision becomes cheap enough to make in places nobody would have paid for one: scoring every passage before it reaches a chatbot, checking every citation a model writes, reading every line of a contract against the standard. Some of the return on decision models will come from work you are already doing. More of it may come from checks you could not previously afford.
How much of your AI work is a decision
No vendor or analyst publishes this number, so we measured the closest thing to it. Anthropic releases a sample of how its business API is used, sorted into task clusters, in its Economic Index data. In the May 2026 data, the clusters whose job is to tag, score, screen, detect, check, route or extract into a fixed structure add up to about 16% of the sampled API conversations. Roughly half of that is closed-set judgment (tagging, scoring, screening, routing), which a decision model can take today. The other half is structured extraction, where a decision model can check a value but not write it.
Other counts come out higher. A June study of 6,003 public automation workflows built in n8n found that the main job of the AI step was extraction in 18% of them and classification in another 10%. Menlo Ventures found that only 16% of enterprise AI deployments are true agents, and that most are fixed-sequence or routing workflows around a single model call.
Our reading: by count of calls in business automation, roughly a sixth to a third of enterprise AI work is decision-shaped. By tokens or spend, the share is smaller, because coding agents and long conversations dominate token volume and a decision call is short. That is why this matters more to latency and to new uses than to the total AI bill. A company whose AI spend is mostly coding assistants will not halve it with a decision model. A company that runs millions of triage, screening and matching calls on a mid-tier model can cut that line by more than 90%.
The fine print
The category is less than a month old, and the evidence behind it is thinner than the marketing.
Accuracy depends on how you ask. An independent benchmark, reported by Beri, took 2,000 emails and asked Jev and Claude Haiku 4.5 whether each one was phishing. Asked once, Jev scored 62.6% and Haiku 81.3%. Split into five narrow questions and combined with a simple fitted formula, Jev reached 95.0%, and a two-line regular expression reached 91.8%. TypeSafe's own guidance says the same thing in its own words: ask "the most explicit, narrow, specific, atomic questions you can." A decision model rewards careful question design and punishes vague questions.
The probabilities are not proven. TypeSafe says Jev is trained to produce calibrated probabilities, but it publishes no calibration measurement, and OpenAI makes no calibration claim at all. Independent tests are mixed. One found Jev's answers at 90% confidence or higher were right 94 to 96% of the time; another found Haiku and Jev tied on accuracy and calibration on a nuanced task. Developers testing OpenAI's version reported choice probabilities pinned at 100% and answers that shifted when the options were reordered, which TypeSafe also lists as a known weakness of Jev. Before you route on a confidence threshold, check that the confidence means what you think it does on your own data.
The headline multiples are not your multiples. TypeSafe's site claims Jev is 193.6 times faster and 444.6 times cheaper, measured against answers produced by two frontier models rather than checked labels, and the company itself says the figures are "on the higher end." An analysis of users' launch-week posts, cited by InfoQ, found a median of 7 times faster and 30 times cheaper. Those are still large numbers. They are just not the ones on the homepage.
Know the limits. Jev reads text only, English best, up to 64,000 tokens per request, and TypeSafe's own list of weak spots includes counting, arithmetic and comparing dates. OpenAI's version accepts images and offers zero data retention, HIPAA eligibility and processing in the US or Europe, but it is in beta, runs on one model, and has no discount for repeated prompts. Anthropic does not have a dedicated decision endpoint: Haiku 5.5 returns the label you ask for through structured outputs, but not a probability for each option unless you ask it to state one.
Your provider, or a new one
The New Stack, covering OpenAI's preview, called the Decisions API "likely a reaction to TypeSafe and Jev." OpenAI has not said so. Either way, a company already on OpenAI or Anthropic now has a cheap decision option from the vendor it already has a contract with, and has to decide whether a new vendor is worth adding for the rest of the saving.
Run the arithmetic before the procurement. The saving from adding Jev is the volume of decisions times the tokens per decision times the price difference. At 700 tokens a call, Jev saves $41 per million decisions against OpenAI's Decisions API. A team making 50 million decisions a month saves about $2,000 a month, which will not pay for a security review, a new data processing agreement and another on-call dependency in the first year. A team scoring 100 million long documents a month at 20,000 tokens each saves more than $100,000 a month, which will.
Price is also not the only reason to add a vendor. Add one when it wins on your own evaluation set: if it is more accurate or better calibrated on the questions you actually ask, or fast enough to fit inside a latency budget your current provider misses. Stay with your provider when you need image input, data residency or contract terms you already have, or when your volume is modest. In either case, keep the decision behind your own interface. The vendors' request formats differ, and the cheapest model this month may not be the cheapest one next quarter.
This article is the argument. The method is in The Decision Layer: Stop Paying Frontier Prices for Yes-or-No Answers, a research paper with a vendor comparison, a cookbook of a dozen decision patterns with example questions and thresholds, an evaluation protocol you can run in two weeks, the break-even arithmetic for adding a provider, and a ninety day plan. Read the executive summary, which is free.
What to do this quarter
For the head of AI platform. Pull a month of calls from your model gateway or logs and sort them by what the model returns: a label, a score, a yes or no, a value picked from a list, or open text. Everything in the first four groups is a candidate. Rank the candidates by volume times the price of the model that serves them today.
For the owner of each candidate workflow. Take 300 to 500 real cases with known correct answers, rewrite the prompt as narrow typed questions, and run them through a decision model alongside the current one. Measure accuracy on the cases the model handles on its own, the share it can handle at a confidence you trust, and the cost and speed of each. Move the volume where quality holds, and keep the large model behind a confidence threshold for the rest.
For finance. Ask for cost per decision, not cost per token, for every high-volume AI workflow, and book the saving when a workflow moves. Then ask what the team would do with decisions at $30 per million, because the largest return here may be in the checks nobody could afford last quarter.
Satori Canton
Founder & Principal
Satori Canton is the founder and principal of ROAI, an advisory practice focused on measuring and improving the return on enterprise AI investment.

