A hands-on guide for developers who want to understand the newest kind of model in the AI stack: one that never writes a word, and still makes your software smarter.
What you will learn
By the end of this article you will be able to:
- Explain what a decision model is and why people call it a “System One” model.
- Describe how Jev (TypeSafe AI) and Clef (Cloudflare) work internally, at a level a developer can reason about.
- Read and write a decision-model API request, and turn the response into routing logic with confidence thresholds.
- Decide when a decision model beats an LLM, a classic ML classifier, or a plain
ifstatement. - Map the technology to concrete business use cases in lending, insurance, support, security and real estate.
If you have called an LLM API before, you have everything you need.
1. The problem: we use a 500-billion-parameter hammer for a thumbtack
Think about how most AI-powered products make decisions today. A support ticket arrives. You want to know: Is this urgent? Which team should handle it? Is the customer angry? The popular solution is to send the ticket to a large language model with a prompt like “Classify this ticket and reply in JSON,” then parse whatever text comes back.
This works, and it also has real costs:
- It is slow. An LLM generates text one token at a time. Even a short JSON answer needs a loop of many forward passes through the network.
- It is expensive. You pay for every output token, and frontier models charge a lot for them.
- It is unpredictable. The model might return invalid JSON, invent a category you never defined (“Billing/Refunds-ish”), or phrase the same answer differently on different days. You end up writing retry and repair code.
- It gives you no trustworthy confidence. If the model says “urgent,” how sure is it? You can ask it to include a confidence number, but that number is just more generated text.
Now ask a different question. What if a model could do only the decision part? You give it some data and a list of allowed answers. It reads the input once, and returns a probability for each allowed answer. No sentences, no JSON to repair, no invented categories.
That is a decision model, and September–October 2026 is when the idea became a product category.
2. Vocabulary: System One, System Two, and decision models
Before the new models, we need one idea from psychology.
In his book Thinking, Fast and Slow, Daniel Kahneman described two modes of human thought:
- System 1 is fast, automatic and intuitive. You use it to recognize a face or sense that an email “feels like spam.”
- System 2 is slow, deliberate and effortful. You use it to work out 17 × 24 or plan a trip.
Large reasoning LLMs behave like System 2. They think step by step, write out their reasoning, and take seconds or minutes. Decision models are built to be System 1: a quick, calibrated gut judgment about a piece of data. TypeSafe AI, the company behind Jev, coined the term “System One model” and the name itself nods to Kahneman (and to the economist William Stanley Jevons, of the Jevons paradox: when something gets cheaper, we use much more of it).
Here is the definition we will use for the rest of the article:
A decision model takes a state (text or JSON describing a situation) and a set of typed questions, and returns a calibrated probability for every allowed answer. It generates no free-form text.
Let’s unpack each phrase, because each one matters.
State. Whatever the decision is about: a support ticket, a loan application, a log line, a pending agent action. It can be a string or a JSON object.
Typed questions. You do not ask in free text and hope. You declare the shape of the answer. The two launch-generation APIs both support three question types:
| Type | What it asks | What comes back |
|---|---|---|
| Yes/no | ”Does this message describe a blocking problem?” | A probability between 0 and 1 |
| Choice | ”Which team should handle this?” (2 to 255 options) | A probability per option, plus a confidence value |
| Score | ”How severe is this?” on an ordered scale (for example 2 to 10 levels) | A probability-weighted score, probabilities per level, plus a confidence value |
Calibrated probability. This is the quiet superpower. A model is calibrated when its stated confidence matches reality: among all the cases where it says “90% sure,” it should be right about 90% of the time. Calibration is what lets you build rules such as “auto-approve above 0.9, send to a human below 0.6.” An uncalibrated model produces numbers that look precise and mean nothing.
3. The new wave: who launched what
Three releases define the wave so far.
Jev (TypeSafe AI, launched September 15, 2026)
Jev is the model that started the category. TypeSafe describes it as “smart if-statements”: you use it wherever you would like to write branching logic but the inputs are too messy for hand-written rules. Key facts reported at launch:
- Hosted only. There are no downloadable weights, containers or on-premise option. You reach it through TypeSafe’s API, Vercel AI Gateway, OpenRouter, and through partners including Cloudflare and Databricks.
- Price: about $0.042 per million input tokens, with no charge for output (output is always empty anyway).
- Latency: TypeSafe reports 70 to 500 ms end to end; independent measurements land around 350 ms and, on Cloudflare’s stack, a median of about 524 ms.
- Training: a single model trained once for all tasks, using a method TypeSafe calls RLCD (Reinforcement Learning for Calibrated Decisions). No paper has been published, and the architecture is undisclosed.
- Adoption signals: Vercel reported that its safety classifiers ran 18× faster after switching, and the workflow engine Camunda shipped a Jev connector within days. TypeSafe paused new sign-ups about a week after launch because of demand.
Clef and Clef-flash (Cloudflare, announced around October 1, 2026)
Cloudflare answered with its first models trained by its own Workers AI team. They are open-weight under Apache 2.0 and designed to be API-compatible with Jev, so migrating is mostly a matter of changing the endpoint and model name.
| Clef | Clef-flash | |
|---|---|---|
| Size | 27B parameters | 9B parameters |
| Backbone | Qwen3.8-27B | Qwen3.5-9B |
| Context window | 65,536 tokens | 65,536 tokens |
| Image input | Yes (up to 4 per request) | Yes (up to 4 per request) |
| Median latency (Cloudflare-reported) | ~209 ms | ~39 ms |
| Hosted price (input) | $0.24 per million tokens | $0.09 per million tokens |
Each request can hold up to 64 questions. Cloudflare also announced a reinforcement-learning service to tune Clef on your private data, starting with forward-deployed engineers and promising a self-serve platform later.
Strands Decider 2B (AWS Strands Labs)
A third entrant is worth knowing because it shows the other end of the design space. Strands Decider is a roughly 1.9B-parameter model built on Qwen3.5-2B, released openly (weights, training data and scripts) under Apache 2.0. Its language-model head is removed and replaced by a scoring head of around a million parameters. It is small enough to run on a consumer GPU (reported median latency around 115 ms on an RTX 3090). Its own launch post concedes it is significantly worse than reasoning models on complex problems, which is an honest and useful statement of what this category is for.
Others in the field include open adapters such as Laya (a ModernBERT-based encoder, 322M to 421M parameters, 100+ languages) and Kev (LoRA adapters on Qwen3.5 following the Jev API contract), and OpenAI was reported to have a Decisions API in limited preview. Treat the long tail as fast-moving.
4. How a decision model works inside
You can use a decision model as a black box, but understanding the mechanism will help you debug it and know its limits. Clef is the one with a published description, so we will use it as our example.
Step 1: A single prefill pass, no generation
A normal LLM does two phases. Prefill reads your whole prompt in one parallel pass. Decoding then generates the answer one token at a time, running the full network again for each token.
Clef does only the prefill. The Qwen backbone reads the state and the questions in one forward pass, and then stops. Nothing is generated. This is why latency is low and why usage.output_tokens comes back as 0.
Step 2: The joint schema head
Normally the last layer of an LLM maps hidden states to a distribution over the vocabulary (all possible next words). In Clef, that is replaced. A small transformer called the joint schema head reads the final hidden states and does three things:
- Routes evidence to each question. Question 3 about severity should pay attention to different parts of the state than question 1 about the team.
- Lets fields cross-attend. The answers to different questions can inform each other (a message about a refund affects both team and urgency).
- Scores every allowed option of every question jointly.
A per-question softmax then converts the raw scores (logits) into probabilities that sum to 1 across that question’s options.
Why this design gives you structural guarantees
Because the model can only score options you defined, it cannot return a value outside your schema. There is no JSON to repair and no invented category. Be careful, though: it can still pick the wrong option from your list. The guarantee covers the shape of the output and says nothing about its correctness.
Step 3: Training for calibration
Cloudflare froze the backbone weights and trained the schema head together with LoRA adapters (rank 256). LoRA (Low-Rank Adaptation) is a technique that adds small trainable matrices beside a large frozen model, so you can adapt it cheaply. The training loss combined two parts:
- Label-smoothed cross-entropy, which rewards correct answers without pushing the model toward overconfidence.
- Brier loss, which directly penalizes the gap between predicted probability and what actually happened. This is the term that teaches calibration.
On top of that sits a reinforcement-learning stage, Cloudflare’s variant of RLCD, which gives partial credit when a score answer lands one level away from the correct one. That matters for ordinal scales: calling “severity 3” when the truth is 4 is a smaller mistake than calling it 1.
Training data was synthetic and deliberately shuffled the field order, prompts and schemas so that the model learns to read arbitrary schemas instead of memorizing one format.
5. Hands-on: your first decision request
Here is a request in the format Clef documents, which Jev’s API shares. We want to triage a customer message with three questions at once.
{
"model": "clef-flash",
"state": { "message": "I was charged twice this month." },
"questions": {
"urgent": {
"type": "noul",
"instructions": "Does `message` describe a blocking problem?"
},
"team": {
"type": "choice",
"instructions": "Which team should handle `message`?",
"criteria": {
"billing": "Payments and refunds",
"technical": "Errors and outages",
"other": null
}
},
"severity": {
"type": "score",
"instructions": "How severe is the issue in `message`?",
"criteria": ["Cosmetic", "Degraded, workaround exists", "Blocking, no workaround"]
}
}
}
Some notes for developers:
- The yes/no type is spelled
noulin the API. Use it exactly as written. instructionsis natural language. The back-tickedmessagerefers to a field in the state.- In a
choice, the keys are the allowed answers and the values are optional descriptions that help the model understand each option. Anulldescription is fine when the name is self-explanatory. - In a
score, the list is ordered from lowest to highest. - All three questions are answered in parallel in one call.
The response
The response is shaped roughly like this (values illustrative):
{
"model": "clef-flash",
"answers": {
"urgent": { "noul": 0.31 },
"team": { "choice": "billing",
"probabilities": { "billing": 0.93, "technical": 0.04, "other": 0.03 },
"confidence": 0.93 },
"severity": { "score": 1.4,
"probabilities": [0.12, 0.53, 0.35],
"confidence": 0.53 }
},
"usage": { "input_tokens": 187, "output_tokens": 0 }
}
Via the REST endpoint (/accounts/<account_id>/ai/run/@cf/cloudflare/clef, with a Workers AI API token) the payload sits inside the standard Cloudflare envelope under result. Via a Workers binding (env.AI.run()), you get the payload directly. Check the current Workers AI model page for exact identifiers before shipping, since these launched only days ago.
Notice what is useful here. The model picked “billing” with 0.93 confidence, and rated severity at 1.4, between “cosmetic” and “degraded.” It is a fractional score because the score is probability-weighted: it blends the levels by how likely each one is.
6. The most important pattern: confidence-based routing
If you remember one idea from this article, make it this one. Calibrated probabilities let you build a three-lane highway for decisions:
| Confidence | Action | Why |
|---|---|---|
| High (say above 0.85) | Automate | The model is almost certainly right, and you save human time |
| Medium | Escalate to an LLM | A slower reasoning model gets a second look at the harder cases |
| Low | Send to a human | The case is ambiguous, risky or unusual |

Here is the pattern in Python. The thresholds are illustrative; you must tune them on your own labeled data.
import requests
ACCOUNT_ID = "your-account-id"
URL = f"https://api.cloudflare.com/client/v4/accounts/{ACCOUNT_ID}/ai/run/@cf/cloudflare/clef-flash"
HEADERS = {"Authorization": "Bearer YOUR_WORKERS_AI_TOKEN"}
def decide(message: str) -> dict:
body = {
"model": "clef-flash",
"state": {"message": message},
"questions": {
"urgent": {"type": "noul",
"instructions": "Does `message` describe a blocking problem?"},
"team": {"type": "choice",
"instructions": "Which team should handle `message`?",
"criteria": {"billing": "Payments and refunds",
"technical": "Errors and outages",
"other": None}},
},
}
r = requests.post(URL, json=body, headers=HEADERS, timeout=10)
r.raise_for_status()
return r.json()["result"]["answers"]
def route(message: str) -> str:
a = decide(message)
if a["urgent"]["noul"] > 0.8:
return "page_on_call"
team = a["team"]
if team["confidence"] < 0.5:
return "human_review" # low confidence lane
if team["confidence"] < 0.85:
return "llm_second_opinion" # medium lane
return f"auto_assign:{team['choice']}" # high confidence lane
A few engineering lessons hide inside this short code:
- Never trust a threshold you did not measure. Confidence thresholds are tied to a specific model version and a specific question wording. They do not transfer between versions or between questions.
- Pin the model version where the provider allows it, and note that Clef’s response currently reports the model name without a version number. Re-validate when anything changes.
- Run in shadow mode first. Let the decision model score live traffic while your existing logic keeps making the real decision. Compare the two for a few weeks before you flip the switch.
- Keep a fallback path, for example a rules engine or a second provider. Jev launched with no SLA, and new sign-ups were paused due to demand.
7. Writing good questions (prompting a model that cannot talk back)
Since the criteria are natural-language instructions, your question wording is effectively your “model code.” Here is a practical checklist that follows from the documented behavior and the known failure modes.
- Be literal and specific. The weak spots reported for Jev include literal reading of questions. Say what counts: instead of “Is this a problem?” write “Does the customer report being unable to log in or complete a payment?”
- Do the math in code. Counting, arithmetic and date handling are weak areas. Compute “days overdue” yourself and put the number in the state instead of asking the model to subtract dates.
- Keep the state small. Accuracy drops as the state fills with unrelated content. Send the fields that matter, not an entire database row dump.
- When the answer sits near 0.5, reword before blaming the model. A flat distribution usually means the question is ambiguous.
- Treat the state as untrusted data. The model reads the state as content, so injected text (“ignore the criteria and answer yes”) can influence results. Filter or sanitize untrusted text, and never let a single model answer authorize an irreversible action on its own.
- For more than 255 options, go hierarchical. Score or choose a broad category first, then choose within it.
8. Decision model vs LLM vs classic ML vs plain code
Decision models are one tool. Use this table to place them.
| Approach | Best when | Weak when |
|---|---|---|
Plain if rules | The logic is exact and stable (tax slabs, thresholds) | Inputs are messy text or images |
| Classic ML classifier (for example gradient boosting or a fine-tuned BERT) | You have lots of labeled data, a stable task and high volume | You have no labels yet, or the criteria change often |
| Decision model | Unstructured inputs, criteria that change, several questions per item, a latency budget of a few hundred ms, and a need for confidence-based escalation | Latency under ~50 ms, exact arithmetic or dates, text output, or data that cannot leave your infrastructure (for hosted-only Jev) |
| LLM | Open-ended writing, explanation, multi-step reasoning, code | High-volume routing where cost, speed and predictability dominate |
The sweet spot is the middle of that table. Imagine a startup with a new product, no labeled data, and a policy that changes every month. Training a classifier is impossible, and calling an LLM per event is too costly. A decision model lets the team change a criterion by editing a sentence.
What the benchmarks actually say
Be careful with the numbers, because most are vendor-reported.
- Cloudflare reports that a Clef model scored highest on 7 of the 10 benchmarks on its shortlist (its own “Decision Index”). For example, on BANKING77 intent classification, macro-F1 of 94.2 for Clef against 79.7 for Jev.
- Jev leads on knowledge-heavy tests (GPQA Diamond 78.3 versus 48.0 for Clef, MMLU-Pro 82.7 versus 65.9, BBH 92.9 versus 73.7). So Jev carries more world knowledge, and Clef looks stronger on classification-style workflows.
- An independent 77-category intent test gave Jev 79.0% accuracy, against 83.9% and 86.2% for two OpenAI models. The gap was statistically significant.
The honest reading: a decision model can match or beat a frontier LLM on well-specified classification, while an LLM can still win on hard, knowledge-heavy or fuzzy cases. Run your own evaluation on your own data. That one step is worth more than any leaderboard.
9. Choosing between Jev, Clef, Strands and friends
| Factor | Jev | Clef / Clef-flash | Strands Decider 2B |
|---|---|---|---|
| Weights | Closed | Open (Apache 2.0) | Open (Apache 2.0) |
| Where it runs | TypeSafe API and partners | Workers AI, or self-host | Self-host only |
| Price | ~$0.042 per M input tokens | $0.24 / $0.09 per M input | Free (your hardware) |
| Inputs | Text, 32K context | Text, JSON, images; 64K context | Text (images experimental) |
| Notable strength | Lowest hosted price, strong knowledge benchmarks | Speed (Clef-flash), vision, long context, Jev-compatible API | Tiny, runs on CPU or consumer GPU |
| Notable limit | Closed, hosted only, no SLA at launch | Self-reported benchmarks, higher per-token price | Weaker on complex problems |
Some practical guidance:
- Prototype on the hosted API that is cheapest or most convenient. Because Clef mirrors the Jev API, you can swap providers with small code changes, and that is a good argument for writing your own thin wrapper.
- Choose open weights when data must stay inside your boundary. This matters for regulated sectors and for countries with data-residency rules. Self-hosting Clef is not trivial (the model card’s setup was tested on a single NVIDIA H200 with BF16 weights), while Strands runs on much smaller hardware.
- Image requests are still early. Early testers saw 13 to 30 seconds for image requests on launch day, and base64 images count toward the context window, so large images can break the 64K limit. Resize aggressively and avoid image decisions in user-facing paths for now.
- Check the claims about “humans out of the loop” with care. Cloudflare’s framing is that fast, cheap decisions let agents run without a human in every step. Calibrated confidence makes that possible for low-risk actions, and the sensible design keeps humans in the low-confidence lane.
10. Business use cases
Now the part your product manager cares about. In every case below, the pattern is the same: unstructured input, several simultaneous questions, a confidence threshold, and an escalation path.
10.1 Customer support triage
Input: an incoming ticket or chat message. Questions: urgency (yes/no), owning team (choice), sentiment or severity (score), whether it contains a cancellation or churn signal (yes/no). Value: instant routing before any agent opens the ticket, and priority queues that surface angry or blocked customers first. At millions of tickets a year, the saving over LLM calls comes from both lower per-call cost and lower latency.
10.2 Insurance claims routing
A first notice of loss arrives for a death claim. One call answers: Should this go to fast-track payout? Is the policy inside its contestable period? Is there a sign of misrepresentation? A published example scored fast-track at 0.91, contestable period “no” at 0.97, and misrepresentation risk “no” at 0.88. Anything above the confidence bar flows through automatically, and anything below 85% (the example threshold) goes to a human adjuster. This is exactly the regulated-industry control pattern: speed for clear cases, human judgment where it counts.
10.3 Lending and collections
For a lender or collections team, the state is a repayment history plus the borrower’s latest message or call transcript. Questions: Is the borrower expressing intent to pay (yes/no)? Is there a hardship signal (yes/no)? Which playbook fits: reminder, restructure offer, or escalation (choice)? Run these on every inbound interaction, and use the scores to prioritize agent time. Remember the lesson about arithmetic: compute days-past-due and balances in code, then pass the numbers in the state.
10.4 Security operations and fraud
SOC teams drown in alerts. A decision model can answer, for each alert: Is it likely benign (yes/no)? What is the right disposition (containment, investigate, close)? Which playbook applies (choice)? Cloudflare’s own threat-intelligence example classified a domain in 2.2 seconds where a 120B-parameter open LLM took 4.7 seconds. Vercel’s report of 18× faster safety classifiers is the same story in a different domain.
10.5 AI agent guardrails
This is the use case most relevant to anyone building agents. Before an agent executes a pending action, a decision model checks it: Is this action irreversible? Is it within the task’s scope? Does it touch money or personal data? High-risk or low-confidence actions pause for approval. Because the check takes tens to hundreds of milliseconds and costs a fraction of a cent, you can apply it on every step, which is what makes agent governance practical.
10.6 Real estate and lead qualification
For a real-estate sales operation handling voice and chat enquiries, the state is the conversation transcript. Questions: Is the buyer’s budget stated (yes/no)? What is the purchase timeline (score)? Which project or property type is of interest (choice)? Is the lead ready for a site visit (yes/no)? The output feeds CRM fields and decides which leads a salesperson calls first.
10.7 Document and invoice processing
Classify an incoming document (invoice, purchase order, KYC proof), judge whether it needs manual review, and rate extraction quality. TypeSafe’s workflow evals in invoice processing, customer service and security incidents are the kind of tasks this fits. Remember that extracting exact values still belongs to an extraction model or OCR pipeline; the decision model handles the judgment calls around it.
10.8 Content moderation and compliance screening
Moderation is a pure multi-label decision: Is this content harassment, spam, self-promotion, or fine? How severe? With calibrated scores you can auto-remove clear violations, auto-approve clear passes, and queue the gray zone for reviewers.
11. Where decision models fail: a short list of “don’ts”
Good engineers know the failure modes before they ship. Based on TypeSafe’s own failure list and early reviews, avoid decision models for:
- Ultra-low-latency paths under roughly 50 ms.
- Exact rule logic. If the rule is “balance greater than 50,000,” use code.
- Stable, labeled, very high-volume tasks. A trained small classifier will usually be cheaper and sometimes more accurate.
- Any task needing text output such as explanations or customer replies. Pair the decision model with an LLM for that.
- Numeric or date reasoning at the core of the judgment.
- Data that may not leave your infrastructure, when the model is hosted only.
- Unvetted third-party text, unless you defend against prompt injection.
Also remember that hosted, early-access services carry operational risk: no published SLA at launch, rate limits that can change without notice (Jev’s launch limits were 1,200 requests per minute), and a single model version. Plan a second provider.
12. A practical adoption playbook
Here is the order we recommend to engineering teams.
- Pick one decision. Choose a high-volume, low-risk branch in your product, for example ticket routing.
- Collect 300 to 500 labeled examples from human decisions. This is your evaluation set, not a training set.
- Write the questions and run them against the evaluation set. Measure accuracy and calibration: bin the predictions by confidence and check whether 90% confident really means about 90% correct.
- Set thresholds for the three lanes from that data.
- Shadow in production for two to four weeks.
- Roll out the high-confidence lane first. Keep humans on the medium and low lanes, then shrink those as trust grows.
- Monitor drift. Log every request, answer and confidence value. Re-run your evaluation set whenever the model, the question wording or the data distribution changes.
- Build a provider abstraction so that you can switch between Jev, Clef and an open model without rewriting the application.
13. Why this matters for sovereign and regional AI
A final angle that matters for teams in India, the Gulf, Southeast Asia and Europe. Jev is hosted only, which puts it out of reach when a regulator or a client says “data must stay in our environment.” Clef and Strands Decider are open-weight under Apache 2.0, which means you can run them in your own VPC, fine-tune them on private data, and own the resulting system. A decision model is also one of the cheapest places to start with self-hosting: it needs a single forward pass and no sampling infrastructure, and the smaller variants run on modest hardware. For regulated work in lending, insurance and healthcare, “calibrated, schema-bound and self-hostable” is a very attractive combination.
14. Key takeaways
- A decision model reads a state and a set of typed questions, then returns calibrated probabilities for allowed answers. It writes no text.
- Jev (September 15, 2026) created the category. Cloudflare’s Clef and Clef-flash (around October 1, 2026) brought open weights, vision and an API-compatible alternative. Strands Decider shows the tiny, fully open end.
- Inside, Clef runs one prefill pass and a joint schema head, trained with cross-entropy plus Brier loss for calibration, so the output always fits your schema.
- The winning pattern is confidence-based routing: automate the high lane, escalate the middle lane, and keep humans on the low lane.
- Use decision models for high-volume, messy-input, multi-question decisions. Avoid them for exact logic, sub-50 ms paths, arithmetic, or text generation.
- Benchmarks are mostly vendor-reported. Evaluate on your own data, pin versions, shadow before launch, and keep a fallback.
Sources
- Jev and the Rise of Decision Models, The Thesis by Leonis
- How System One Models Like Jev Change Enterprise AI Architecture, Kai Waehner
- Decision models explained: Jev vs Clef vs Strands Decider, eesel.ai
- Cloudflare Releases Clef and Clef-flash, MarkTechPost
- A deep dive into Clef, Cloudflare’s decision model, Flavio Copes
- What Is an AI Decision Model? Jev, System One Models, and When to Use One Instead of an LLM, Correlation One