OpenAI put the Decisions API into public beta on October 6, a week after announcing it at DevDay as a limited preview. It is a separate endpoint, POST /v1/decisions, that runs only on gpt-6-luna and returns typed answers rather than text: a probability that a condition holds, one choice from a list you supply, or a score against ordered levels. OpenAI says it expects general availability “in the coming weeks”. Decisions guide API changelog Our DevDay article listed the missing model ID, endpoint, schema and price as items to watch; the guide now supplies all four.
The price is the part that changes a cost model. Input is $0.10 per million tokens, the same as Luna on the Responses API, and there is no charge at all for output, cache reads or cache writes. For work where the model only has to pick, that removes the output side of the bill. It does not make every classifier cheaper, and the cases where it loses are worth knowing before anything gets rewritten. S5 Labs has not run the endpoint; the speed and pricing figures below are OpenAI’s, and the cost examples are arithmetic on list rates.
What a decision call looks like
A request has three parts: the model, an input that is either a string or user messages mixing text and images, and a questions array. Each question has a type, a name, instructions, and for two of the types a list of allowed answers. The response is an answers array keyed by the same names, so one request can carry several independent questions over the same input. Decisions guide
| Type | What it returns | Use it for |
|---|---|---|
predicate | probability, 0 to 1, that the condition is true | Damage visible in a photo, passage relevant to a query, message is a cancellation |
choice | choice from your values, plus probabilities over all of them and a confidence | Department routing, content category, document type |
score | score, the probability-weighted average of the level indices, plus probabilities and confidence | Severity, urgency, any ordered rubric |
OpenAI’s routing example sends a one-line complaint and four departments:
curl https://api.openai.com/v1/decisions \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-6-luna",
"input": "I was charged twice for my order.",
"questions": [{
"type": "choice",
"name": "department",
"instructions": "Which department should handle this complaint?",
"choices": [
{"value": "billing", "description": "Payments, invoices, and refunds."},
{"value": "technical", "description": "Problems using the product."},
{"value": "shipping", "description": "Delivery and tracking."},
{"value": "other", "description": "Requests outside these categories."}
]
}]
}'
The illustrative answer is billing with a confidence of 0.93. The score type behaves differently from a plain label: in OpenAI’s severity example, probabilities of 0.1, 0.7 and 0.2 across cosmetic, workaround available and fully blocked produce a score of 1.1, between the second and third level. That is the point of the type: a number you can sort and threshold rather than a bucket, with the per-level probabilities showing how spread the model’s view was. Decisions guide
Four rules from the guide shape how you design around it. Questions that depend on an earlier answer need a separate request, so a two-stage triage (is this damaged, then which repair category) is two calls. Images must be inline base64 data URLs; hosted URLs and file_id references are rejected, so you re-send the bytes every time rather than referencing an upload. OpenAI recommends a fallback option such as “other” whenever your categories might not cover every input. And the SDK examples check for an answer of type refusal before reading a result, so a classifier built on this needs a branch for the model declining to answer. Decisions guide
The guide draws its own boundary with the rest of the platform. Use Decisions when the application needs one of the three answer types. Use Structured Outputs on the Responses API when you need an object that follows your own JSON schema, such as extracted fields or a written explanation, and function calling when the model should request a tool call with arguments. Decisions guide In practice that means Decisions replaces the enum-only slice of structured output work. Anything that also pulls out a name, a date or a reason stays where it is.
The cost comparison
The guide’s pricing paragraph is short: with gpt-6-luna, input costs $0.10 per million tokens, and “You pay only for input tokens: there are no cache-read, cache-write, or output-token charges.” Regional processing premiums and long-context input multipliers still apply. Decisions guide From the pricing and data-controls pages, those are a 10% uplift on regional-processing (data-residency) endpoints for models released on or after March 5, 2026, and doubled input and cache rates on prompts above 272,000 input tokens. Pricing Data controls GPT-6 Luna model page Neither page says outright whether Luna falls under the 10% uplift, so the data-residency figures below assume it does. The pricing page itself has no Decisions row yet, and the Luna model page still lists Chat Completions, Responses and Batch as its endpoints without /v1/decisions; the guide is the only place the price appears.
Take a worked example: one million support tickets, 500 input tokens each, routed to a department. That is 500 million input tokens.
| Route | Input | Output | Total |
|---|---|---|---|
| Decisions, Luna | $50 | $0 | $50 |
| Decisions, data-residency endpoint (assumes the 10% uplift applies) | $55 | $0 | $55 |
Responses + structured output, Luna, reasoning.effort: none, ~20 output tokens | $50 | $10 | $60 |
Responses + structured output, Luna, default medium effort, 200 reasoning tokens assumed | $50 | $110 | $160 |
Input is identical on both routes because the rate is the same $0.10. The difference is entirely the output side. A minimal JSON answer like {"department":"billing"} is about 20 tokens, which costs $10 per million tickets at Luna’s $0.50 output rate. Reasoning tokens are also billed as output tokens, and Luna defaults to medium effort, so a Responses call that is allowed to think before it answers can multiply that line. Reasoning guide GPT-6 Luna model page The 200-token reasoning figure in the last row is an assumption to show the shape of the cost, not a measurement; your own usage logs will give the real number. The guide does not say what effort, if any, the Decisions endpoint applies internally, and documents no parameter to set it.
On a short-input, label-only job, then, Decisions saves the cost of the output tokens: about a sixth of the bill against a Responses call with reasoning off, and much more against one with reasoning on. That is real money at scale and close to nothing at a few thousand calls a day.
The comparison turns over when the shared context is large. Decisions has no cache, so every token in every request is billed at $0.10. The Responses API bills cached input to Luna at $0.01. Pricing Put a 4,000-token policy document in front of each 500-token ticket:
| Route | Shared prefix | Ticket | Output | Total per 1M tickets |
|---|---|---|---|---|
| Decisions | 4,000M × $0.10 = $400 | $50 | $0 | $450 |
| Responses, every call a cache hit, effort none | 4,000M × $0.01 = $40 | $50 | $10 | $100 |
That assumes a cache hit on every call and ignores the one-off cache write, so it is the best case for Responses. The arithmetic favors Responses from a prefix of little more than a hundred tokens, but caching only applies from 1,024 tokens on GPT-5.6 and later, and reuse needs the entire rendered prefix to match, so the practical crossover sits around that minimum. Prompt caching guide Decisions is the cheaper route when the input is mostly the thing being judged. When the input is mostly a reusable rulebook, the Responses API with caching wins, and a long rulebook is how many moderation and compliance prompts are written. The guide does not say whether the questions block, with its choice descriptions, is metered as input; assume it is, and keep the descriptions short.
The speed claim and its basis
The guide’s first sentence is the claim: the Decisions API “returns typed answers about 10x faster than the Responses API”, and the changelog entry repeats “10x faster than the Responses API”. Decisions guide API changelog Neither page says which Responses configuration it is measured against, at what effort setting or input size, or whether the comparison is time to first token or time to a complete parsed answer. Our DevDay article relayed a secondhand version, ten times faster than Luna through the regular API, from The Decoder; the published docs compare against the Responses API, and that is the wording to use.
The mechanism is plausible, since a Responses call generates the JSON token by token, possibly after hidden reasoning tokens, while a Decisions call reads the input and emits a distribution over a fixed set. A plausible mechanism is still not a measurement. For a routing step inside a latency budget, run both against the same hundred inputs and record p50 and p95 end to end, with the Responses side at none effort as well as the default, so the comparison is against the fastest configuration you would have used.
What the probabilities are good for
The guide is careful about what the probabilities are. A predicate’s probability is “the model’s estimate that the condition is true”; choice and score answers come with a distribution and a separate confidence. It makes no calibration claim. Its advice is to use labeled examples from your own application to set thresholds for routing, filtering or review, chosen on the cost of false positives against false negatives. Decisions guide
That is the right advice, and it is more work than the JSON-mode version of the same classifier, where the output is a label and the threshold is implicit. Treat the number as a ranking signal until you have checked it: run a few hundred historical tickets with known departments and watch accuracy as you move the confidence cutoff, routing everything below it to a person. If the probabilities are reasonably calibrated on your data, the score type gives you a severity queue for free. If not, you still have a usable classifier with an empirically set cutoff.
Two design notes follow from the guide’s own rules. Keep one concern per question: a damage check and a product-category check belong in the same request as two questions, but “is it damaged and is it a return” as one predicate gives you a probability of an ambiguous conjunction. And define adjacent score levels so that they have distinct criteria, because the score is an average across them and blurry levels produce a blurry number.
What else changed on the same day
The October 6 changelog carries a second entry: API usage tiers went from five to three, named Build, Launch and Grow, with an organization’s tier upgrading automatically as its cumulative credit purchases reach each threshold. API changelog Rate limits guide The rate-limits page has the thresholds and per-model limits; they are the kind of figure that moves, so check that page rather than this one. The Decisions endpoint is not mentioned there, and the guide publishes no rate limit for it.
A day earlier, October 5, OpenAI added a self-serve flow under API Organization settings > General where admins of eligible organizations can accept the standard Business Associate Agreement. API changelog That matters here because the data-controls page lists /v1/decisions as eligible for HIPAA use under an executed BAA and healthcare addendum, as ZDR-eligible with a caveat about prompt caching state, and as supported for US and EU residency and regional processing. Data controls A healthcare intake triage on this endpoint is possible on paper from day one of the beta; any residency uplift is the 10% above.
The open-weight alternatives
Cloudflare’s Clef, a 27B decision model on Workers AI, takes the same shape of request, a state plus typed questions, and returns a probability per allowed answer, at $0.24 per million input tokens with a 65,536-token context and no output price listed. Clef on Workers AI Liquid AI published d1-3B, an open-weight decision model that answers “in a single forward pass” rather than producing tokens, on October 7. Liquid AI on Hugging Face Those run where you choose, which the Decisions API cannot, and they come with published latency and benchmark figures of their own. Our open decision models comparison sets Clef, Clef-flash, Liquid d1 and TypeSafe’s Jev against this endpoint on license, price, context and input types.
What to check before moving a classifier
Start with the shape of the prompt, since that decides the bill. If the per-item input outweighs the shared instructions, Decisions is cheaper on list price. If the instructions outweigh the item and you already get cache hits on the Responses API, it is probably not; a $0.01 cached rate is hard to beat with a $0.10 uncached one.
Then check what the step actually returns. If downstream code consumes only a label or a number, the move is clean. If it also consumes a field the model extracted or a sentence it wrote, you would be paying for two calls where one worked, and the guide itself points you back to Structured Outputs for that.
Last, test the probabilities on your own labeled data before trusting a threshold, time both routes on the same inputs rather than taking 10x on faith, and note the beta constraints: one model, base64 images only, no cache, no stated rate limit, and GA “in the coming weeks”. The endpoint fits high-volume routing, moderation, lead scoring and relevance filtering where the input is short and the answer is a pick. It fits less well than it first looks when a long rulebook sits in front of every call.
Key details
| Item | Detail |
|---|---|
| Endpoint | POST /v1/decisions, public beta from October 6, 2026; GA expected “in the coming weeks” |
| Model | gpt-6-luna only |
| Question types | predicate (probability), choice (choice, probabilities, confidence), score (weighted average of level indices, probabilities, confidence) |
| Input | Text, or user messages with text and base64 image data URLs; no hosted image URLs, no file_id |
| Price | $0.10 per million input tokens; no output, cache-read or cache-write charges; 10% data-residency uplift for models released on or after March 5, 2026 (not confirmed for Luna); long-context multiplier above 272K input tokens |
| Speed | OpenAI: “about 10x faster than the Responses API”; no method published |
| Data controls | ZDR and HIPAA for eligible customers; US and EU (EEA + Switzerland) residency |
| Same-day changes | Usage tiers reduced to Build, Launch and Grow (Oct 6); self-serve BAA acceptance in Organization settings (Oct 5) |
| SDK minimums | Python 3.26.0, JavaScript 7.30.0, Go 3.73.0, Ruby 0.101.0, Java 4.78.0 |
Sources
- Decisions API guide — OpenAI Developers
- API changelog — OpenAI Developers
- Pricing — OpenAI Developers
- GPT-6 Luna model page — OpenAI Developers
- Reasoning guide — OpenAI Developers
- Rate limits and usage tiers — OpenAI Developers
- Data controls — OpenAI Developers
- Clef model page — Cloudflare Workers AI
