Decision Models Go Open Weight: Cloudflare Clef, Liquid d1 and TypeSafe Jev Compared

Cloudflare's Clef and Liquid's d1-3B are open-weight decision models. Compared with Jev and OpenAI's Decisions API on license, price, context and limits.

A decision model takes a piece of state and a set of typed questions, and returns a probability for every allowed answer. It writes no text. TypeSafe sells that idea as a hosted API called Jev, and OpenAI opened its own Decisions API to public beta on October 6 with the same shape. The change this week is that the weights are now public: Cloudflare published Clef and Clef-flash on October 1 under Apache 2.0, and Liquid AI published d1-3B on October 7 under its own LFM Open License. Cloudflare, Clef announcement Liquid AI, Open d1 d1-3B model card For a team that runs classification, routing or triage through an LLM in JSON mode, there are now four hosted endpoints priced per input token and three models you can run on your own hardware. S5 Labs has not run any of them. Every benchmark and latency figure below is the vendor’s.

The OpenAI endpoint has its own article, OpenAI Decisions API public beta, which covers the pricing, question types and worked costs. This piece is about the open alternatives and when any of the four should replace a generation call.

What a decision model replaces

The common pattern today: send a support ticket to a chat model with a system prompt that says “answer with JSON matching this schema”, parse the result, and retry when the JSON is malformed or the label is outside the allowed set. The model generates the answer token by token, so you pay output tokens for the JSON and wait for them. The “confidence” field, if you asked for one, is a number the model wrote rather than anything measured.

A decision model runs one forward pass over the input and scores the allowed options directly. Cloudflare describes Clef’s decision step as non-autoregressive, with no intermediate text to generate. Cloudflare, Clef announcement Liquid says the same of d1: the models “don’t produce tokens but answer in a single forward pass”. Liquid AI, Open d1 The output is a probability distribution over the options you defined, so there is nothing to parse and no label outside the set.

All four products expose the same three primitives, with one naming difference. A yes/no question returns a probability (Jev, Clef and d1 call it noul, OpenAI calls it predicate). A choice picks one option from a labelled set and returns per-option probabilities plus a confidence. A score places the input on an ordered scale and returns a probability-weighted position that can fall between levels. TypeSafe, API reference Cloudflare, Clef model page Liquid AI, d1-3B model card OpenAI, Decisions guide Cloudflare describes Clef as fully compatible with the Jev API, and the request shape on Jev, Clef and d1, state plus a questions map, is close enough that switching between them is mostly a URL change. Cloudflare, Clef announcement

None of them produces a value you did not enumerate. A decision model can tell you a ticket is about billing with probability 0.91. It cannot extract the invoice number, summarize the complaint or draft the reply. Jev’s score type takes 2 to 10 levels and its choice type up to 255 options; the same 2-to-10 range for scores is on the d1-3B card. TypeSafe, API reference Liquid AI, d1-3B model card If the JSON your LLM returns contains any free-text or open-numeric field, a decision model replaces only the part of it that is a bounded choice. A common split is a decision model for the gate (is this spam, which queue, how urgent) and a generation call only for the fraction of traffic that passes it.

The five options side by side

Jev 1.13ClefClef-flashd1-3BOpenAI Decisions
VendorTypeSafeCloudflareCloudflareLiquid AIOpenAI
Base modelNot disclosedQwen3.8-27B, frozenQwen3.5-9BLFM2.5-VL-3BGPT-6 Luna
ParametersNot disclosed27B9B3.12BNot disclosed
LicenseHosted service; no weightsApache 2.0Apache 2.0LFM Open License v1.0Hosted service
Self-hostableNoYesYesYesNo
Hosted price$0.042 per 1M input tokens, output free$0.24 per 1M input tokens (Workers AI)$0.09 per 1M input tokens (Workers AI)No hosted endpoint in Liquid’s post$0.10 per 1M input tokens; no output or cache charge
Context64K per request; 32K for state plus the longest question65,536 (Workers AI)65,536 (Workers AI)32,768Not stated in the guide; Luna’s model page lists 1,050,000
InputsText and JSON onlyText, JSON, up to 4 inline imagesText, JSON, up to 4 inline imagesText, JSON, imagesText and base64 images
Questions per requestNot stated1 to 641 to 64Not statedNot stated
StatusCurrent (jev-latest = 1.13.0)Released Oct 1Released Oct 1Released Oct 7Beta, GA “in the coming weeks”

Sources: TypeSafe models and pricing, Clef model page, Clef-flash model page, Workers AI pricing, Clef on Hugging Face, Clef-flash on Hugging Face, d1-3B model card, OpenAI Decisions guide, GPT-6 Luna model page.

Two notes on the price column. Cloudflare’s model pages and its Workers AI pricing table list only an input price for Clef and Clef-flash, with no output row; that is not the same as a stated zero, so we have not written one. OpenAI’s guide does say output, cache reads and cache writes are not charged on /v1/decisions, but adds that regional processing premiums and long-context multipliers apply, and the general pricing page has no Decisions row at all. OpenAI, Decisions guide OpenAI pricing

TypeSafe’s context figure needs the footnote. Its models page gives 64K tokens per request for the state plus all questions combined, and 32K for the state plus the single longest question, which is why Cloudflare’s blog lists Jev at 32K. TypeSafe models and pricing

Where each one falls down

Jev is text only. TypeSafe’s docs say images, audio and video are not supported (yet), so anything visual has to be described in text first. TypeSafe, System One The rate limits are 100K tokens per second and 80 requests per second, and the models page warns that they are changing dynamically and may change without notice. TypeSafe models and pricing There are no weights, so there is no fallback if the service or the pricing changes. In exchange it is the cheapest hosted option by a wide margin, and on Cloudflare’s own tables it still leads on the reasoning-heavy rows: GPQA Diamond 78.3 against Clef’s 48.0, MMLU-Pro 82.7 against 65.9, BBH 92.9 against 73.7. Clef on Hugging Face If your questions need the model to reason rather than recognize, that gap matters more than the intent headline.

Clef is the expensive one on Workers AI, at $0.24 per million input tokens, more than twice OpenAI’s rate and almost six times Jev’s. It earns that on Cloudflare’s intent benchmarks: CLINC150+OOS macro-F1 of 97.43 against Jev’s 89.27, BANKING77 94.2 against 79.7. Cloudflare, Clef announcement Self-hosting is where the license pays off, and where the bill moves to hardware. Clef is a frozen Qwen3.8-27B with rank-256 low-rank adapters and a jointly trained head, so it is a 27B-parameter model. Cloudflare, Clef announcement Cloudflare’s tested setup is a single H200, and the base checkpoint in BF16 is 55.6 GB per our Qwen3.8-27B article. Clef on Hugging Face The card links community quantizations for llama.cpp, Ollama and LM Studio, which we have not tried. The weights are tagged custom-code, so loading them means trusting Cloudflare’s joint_schema_model.py, and the card’s encode_record helper defaults to a 16,384-token maximum, against the 65,536 the hosted endpoint advertises. We have not tested whether the self-hosted model handles the longer input. The LICENSE file in the repository is the standard Apache 2.0 text, with Alibaba Cloud named as copyright holder, the base model’s notice rather than Cloudflare’s. Clef LICENSE

Clef-flash is the one Cloudflare pitches for latency, and it wins the tool-calling rows (BFCL 98.76, API-Bank 93.11). But the same table has it at 66.77 on CLINC150+OOS, about thirty points under Clef and twenty-two under Jev, and at 35.6 on RAGTruth hallucination F1 against Clef’s 79.4 (the RAGTruth figures are on the Hugging Face card only). Cloudflare, Clef announcement Clef-flash on Hugging Face An intent classifier with 150 labels is exactly the job a flash model is sold for, so the drop matters. Test it on your label set before taking the $0.09 rate.

d1-3B is the smallest and the only one with a license you have to read. The LFM Open License v1.0 grants a royalty-free, irrevocable license to use, modify and distribute, but Section 5 ties commercial use to a threshold of “annual revenue of 10 million United States dollars ($10,000,000) or more”: commercial use by a legal entity above it “is not licensed under this Agreement”. Qualified non-profits are exempt for research and non-commercial use. d1-3B LICENSE The license defines the licensee to include entities under common control, and does not say how revenue is counted across affiliates or clients, so an agency shipping d1 in client work should read Section 5 against both its own revenue and the client’s, and talk to Liquid if either is near the line. The license terminates automatically on any breach. Context is 32,768 tokens, half of Clef’s. Liquid’s post describes no hosted endpoint, so this is a run-it-yourself model, and Liquid’s numbers are built around that: 8 ms for a single warm question on an RTX 4090, 30 ms on an Apple M5 Pro, 50 ms on a Jetson Orin Nano, all vendor-measured. Liquid AI, Open d1 On quality, Liquid reports a Decision Index 0.2.1 score of 48.57, ahead of Decider 35B-A3B at 47.11, and calls it the best decision model under 10B. Its comparisons are against the Decider models only; it publishes nothing against Clef or Jev, and its blog and model card quote different benchmark means for d1-3B (82.9 and 77.1, the latter across eight internal benchmarks); we could not tell how the two sets differ. Liquid AI, Open d1 Liquid AI, d1-3B model card

OpenAI Decisions runs on GPT-6 Luna only, is in beta, and the guide publishes no context limit, rate limit, or maximum number of questions, choices or levels. Images must be inline base64 data URLs; hosted URLs and file_id are rejected. Decisions that depend on each other need separate requests. OpenAI, Decisions guide The attraction is that it sits in an account you already have, with ZDR, HIPAA use and US or European data residency for eligible customers, at a price between Clef-flash and Clef.

Latency: three sets of numbers that do not compare

Every vendor leads with speed, and each measured something different.

Cloudflare reports median latency across its 43 evaluation benchmarks: Clef 209.3 ms, Clef-flash 38.8 ms, Jev 524.1 ms, with p95 at 238.6, 122.4 and 536.0 ms. It does not describe the hardware or where the clock starts. Jev has no public weights, so its figure can only be Cloudflare timing TypeSafe’s hosted API from wherever Cloudflare ran the tests, which the post does not spell out. Cloudflare, Clef announcement Liquid’s 8 ms is a single warm question on a local RTX 4090 with no network in the path. Liquid AI, Open d1 OpenAI’s claim is relative: “about 10x faster than the Responses API”, with no absolute figure and no method. OpenAI, Decisions guide

Putting 209 ms next to 8 ms says nothing about Clef against d1; it says a 27B model behind a network is slower than a 3B model on the card in front of you. The comparison worth making is the one you can run yourself, with your own state size, question count and deployment. Liquid’s table shows how fast the small model’s numbers move with input, from 8 ms for one question to 102 ms for a 3.4K-token state on the same 4090, so a ticket with a long thread attached will not see the headline figure either. Liquid AI, Open d1

What a million tickets cost

Take a round example: 1 million tickets at 500 tokens each, so 500 million input tokens, hosted list prices, no regional premium.

RouteInput costOutput costTotal
Jev 1.13$21.00$0$21
Clef-flash on Workers AI$45.00Not listed$45 plus any output charge
OpenAI Decisions (Luna)$50.00$0$50
Clef on Workers AI$120.00Not listed$120 plus any output charge
GPT-6 Luna, Responses API, JSON output of 40 tokens$50.00$20.00$70

The Luna row uses the pricing page’s $0.10 input and $0.50 output rates. OpenAI pricing Luna in JSON mode is already cheap, and the output tokens for a short label add 40% to it. The gap opens against a mid-tier generative model: our GPT-6 Sol and Luna article lists Sol at $2 per million input and $10 per million output, so the same job on Sol is $1,000 of input plus $400 for the same 40-token labels. If you route through Sol-class or larger models because the small ones mislabel, a decision model is cheaper by a factor of about 12 (Clef) to 67 (Jev) at list prices. Whether it labels as well is a separate question; Cloudflare’s and Liquid’s tables give you a reason to run your own test and do not settle it.

For the self-hosted rows the token price is zero and the cost is a GPU. d1-3B on a workstation card or a Jetson is the realistic low end; Clef needs a datacenter-class card at full precision. A team already running Qwen3.8-27B quantized on a 24 GB card is the natural Clef self-host candidate, if the community quantizations of the adapted model hold up; the Qwen3.8-27B article covers what that card can hold.

When to make the switch

Replace the JSON-mode call when all of the following are true: every field in the schema is a bounded choice, a yes/no or a rating; the labels are stable enough to train or prompt against; and you want a probability you can threshold rather than a label you have to trust. Routing, spam and abuse gates, urgency scoring, lead qualification and “did the agent finish the task” checks fit. Extraction, summarization and anything with a free-text field do not, and a schema that mixes the two should be split, with the decision model in front.

Which of the four depends on constraints you already know. Text-only traffic at the lowest hosted price points to Jev, with no exit if TypeSafe changes terms. Image inputs on a hosted API means Clef, Clef-flash or OpenAI. Keeping data on your own hardware leaves Clef and d1, and the choice between them is a license check and a GPU check: d1 if you are under $10 million in revenue or willing to negotiate, and have a modest card; Clef if you are over the threshold and have the memory. Teams already inside OpenAI’s compliance program will find the Decisions API the shortest path, and should read the beta’s missing limits as a reason to keep the Responses fallback wired up.

Whichever route, the first test is the same: take a week of real traffic with the labels you already have, run it through the candidate with your question schema, and plot the probabilities against the outcomes. Calibration is what these models claim to sell. TypeSafe’s docs say calibration is measured across groups of predictions rather than per answer, and Cloudflare trains Clef with a Brier loss for the same reason. TypeSafe, System One Cloudflare, Clef announcement A model whose 0.9 is right nine times in ten lets you automate the confident cases and route the rest to a person. A model whose 0.9 is right six times in ten is a classifier with a decorative number on it, and no vendor table will tell you which one you have.

Continue reading.

Insight14 min read

Claude Haiku 5.5: A 90% List-Price Cut, a 75% Real One, and Five Ways to Get a 400

Haiku 5.5 lists at a tenth of Haiku 4.5's price; Anthropic says about 75% less in practice. The gap, the 400s after switching, and Sonnet 5.5 vs Luna.

Insight13 min read

OpenAI Decisions API: $0.10 per Million Input Tokens and Nothing for the Answer

OpenAI's Decisions API: public beta, GPT-6 Luna only, $0.10 per million input tokens, no output charge. Where it beats Responses with structured output.

Insight3 min read

Claude Sonnet 5.5: Compare Cost per Finished Task

Sonnet 5.5 keeps Sonnet 5 token prices. Its claimed savings come from using fewer tokens. How to test that claim in your workflow.