AI decision models: what they are and where they fit

ARTIFICIAL INTELLIGENCE, BLOG.AUTOMATIZACIÓN.
Digital illustration: lines of glowing dots pass through a narrow teal gate on a dark background and leave as three parallel tracks.

In short: AI decision models answer narrow questions from predefined options and return a probability software can act on. Their advantages are speed and cost. In independent tests, Jev, TypeSafe AI's decision model, answered in a median of 148 ms at $0.11 per 1,000 requests, against 3,609 ms and $1.18 for a general large language model (LLM) measured in batches on one GPU. Since Jev's release on 15 September, the category has grown into a competitive market. Once labeled data exists, small trained classifiers can match or outperform them, and the wording of each question changes the answers. A decision model is one stage in the lifecycle of a decision, alongside rules, trained classifiers and LLMs.

This guide is for organizations whose software routes requests, classifies documents or flags cases for human review. It covers the speed and cost of AI decision models, when a trained classifier remains the better option, and how a recurring decision can move from one approach to another as it matures. One note on terminology: in process management, "decision model" already refers to a Decision Model and Notation (DMN) model, decision logic written as tables a person can read rule by rule. In this guide, "decision model" means the AI kind; the final section explains how they relate.

What an AI decision model is and is not

An AI decision model is a model built to make one type of judgment inside software. It receives the facts of a case and one or more questions whose possible answers are defined in advance. For each question, it returns a typed value: a probability for each option, a position on a scale or the probability of a yes. It does not generate text or explain its reasoning. Its output is an answer that software can store, compare against a threshold and act on, rather than text that has to be interpreted. This makes decision models suited to the narrow, repeated judgments that business processes make at high volume, such as routing a request, classifying a document or deciding whether a case needs human review.

TypeSafe AI, which released Jev on 15 September 2026, calls these System One models. The name comes from Daniel Kahneman, a Nobel laureate in economics, who called fast, automatic thinking System 1, and slow, effortful thinking System 2, in Thinking, Fast and Slow, his 2011 book on how people think and make choices.

Consider a customer onboarding application. The software sends the model typed questions defined in code and a state: the facts of the case, here the product requested, the documents received and the applicant's note. The model returns one value per question.

Decision models support these question types:

  • Choice: selects one of up to 255 predefined options, each named and optionally described in plain language, with a probability for each.

  • Score: places the case on a scale of 2 to 10 levels described in the question.

  • Yes/no: returns a single probability. The question can describe what yes and no mean. Jev's API calls this type "noul".

The probability is designed to be calibrated: outcomes assigned a probability of 0.2 should occur about 20% of the time. TypeSafe presents this as a property of many answers taken together, not a promise about any single answer, and its documentation does not publish a measured figure.

A decision model returns typed values, not text: a next step with a probability per option, a completeness score and a yes/no probability, with illustrative values
A decision model returns typed values, not text: a next step with a probability per option, a completeness score and a yes/no probability, with illustrative values

The term System One describes what a model returns, not how it is built, and other providers have since released their own decision models. AWS's Strands Decider 2B is a 2-billion-parameter language model whose text output is replaced by a final layer that rates each option. Convai Innovations' Laya builds its English checkpoint, a trained version of the model, on ModernBERT, a much smaller model of 421 million parameters.

The consultancy HatchWorks distinguishes a valid answer, a correct answer and an authorized action, which are easily confused. By design, a decision model can only return one of the options it was given, so its answers are valid. An evaluation on the organization's own data shows whether they are correct, and the organization's policies and people decide whether the system may act on them.

A decision model also differs from a trained classifier. A trained classifier learns its categories from one task's labels, that is, past cases marked with the correct answer, and needs retraining when those categories change. A decision model receives its options in plain language with each request and can answer several questions per call.

Decision model vs LLM: speed, cost and output

Decision models have a clear advantage in speed and cost. Independent tests by researchers at Texas State University, published as a preprint and not yet peer-reviewed, measured Jev and a general LLM on intent classification, which sorts requests by what the user wants. When the LLM's answer was read from the probability it assigned to each option, it matched Jev's accuracy: 0.912 against 0.913. The LLM took considerably longer and cost more per request. The table compares a decision model with an LLM:

AI decision model

Large language model

What it can answer

Only the questions and options defined in advance

Open questions, plans and text

What it returns

One typed value per question, from the defined options

Text that the application must parse, even in JSON mode, which constrains the format but still returns text

Confidence

A probability per option, designed to be calibrated across many answers

Stated in its reply when requested, or read from the probability it assigns to each option

Explanation

None: no text and no rule to point to

In words

Speed in one test

148 ms median, Jev through its API

3,609 ms median, a 27-billion-parameter model measured in batches on one A100 GPU

Cost in one test

$0.11 per 1,000 requests at Jev's list price

$1.18 per 1,000 requests at an assumed hourly rate for that GPU

Speed and cost figures from those tests. The LLM's figures are batch-amortized: each batch's time and cost is spread across its requests. Each side was measured differently, so the figures should not be divided into a ratio.

TypeSafe advertises Jev as 193.6 times faster and 444.6 times cheaper than LLMs on its own workflows. It does not name the models it compared against, runs the LLMs through its own wrapper and expects these figures to sit at the upper end of real-world gains. Kai Waehner, an independent enterprise architect, expects cost savings of about one order of magnitude across a complete system.

Which AI decision models are available, and how to access them

Decision models released by early October, each linked to its provider's page:

Model

Provider

Released

Access

What stands out

Jev

TypeSafe AI

15 Sep

Early access, hosted API

$0.042 per million input tokens, output free; 70–500 ms measured by TypeSafe from the US West Coast

Kev

Jared Palmer, formerly vice president of AI at Vercel

17 Sep; Kev 1.0 on 1 Oct

Apache 2.0 weights, 0.8B to 27B parameters

Accepts Jev's request format

Laya

Convai Innovations

18 Sep, repo created

Apache 2.0 weights; self-hosted

English and multilingual checkpoints; 33 to 40 ms per question on a Tesla T4 GPU; built to be fine-tuned per task

Decisions API

OpenAI

29 Sep

Limited preview

Runs on OpenAI's Luna model; text or images as context; no public documentation or pricing in its first days

Strands Decider 2B

AWS Strands Labs

1 Oct

Downloadable, with training data and scripts

A median of about 115 ms on an RTX 3090 GPU

Clef, Clef-flash

Cloudflare

1 Oct

Apache 2.0 weights; hosted on Cloudflare's Workers AI

Accepts images; accepts Jev's request format

A community directory lists more than a dozen. All figures are the providers' own.

Kai Waehner reports that several platforms added Jev access by 25 September. Availability, however, does not show whether a model's answers are correct for a given business process.

What determines results in practice: labeled data, calibration, question design and cost

These factors determine how well a decision model performs in a business setting, according to the same independent tests of decision models including Jev, Laya and Kev; not every model in the table above was part of them. The figures refer mainly to Jev:

  1. Labeled history. Once a task has labeled history, small trained classifiers can be more accurate: 0.948 and 0.940 on intent classification, compared with 0.913 for Jev. On business-workflow decisions, the difference was not significant.

  2. Calibration. Probabilities that were reliable on one task were not reliable on another, so each task needs its own thresholds.

  3. Question design. Software acts on an answer only when its probability exceeds a threshold. Out-of-scope requests are those that fit none of the options. With the threshold set for 5% errors on in-scope requests, Jev still accepted 31% of out-of-scope requests (reported uncertainty range: 14% to 50%), assigning each to an option with a probability above the threshold. Adding an explicit "none of these" option reduced that figure to 19.5%. On business-workflow questions, swapping the names of the yes and no options so that each contradicted its written description changed about half of the answers.

  4. Cost structure. A cascade can lower cost: a small decision model fine-tuned on the task's labeled data (not a classifier) passed only its uncertain cases to Jev, and the cascade matched Jev's accuracy at 43% of the cost of running Jev alone. This result applies to intent classification only and assumes fully used GPUs.

Labeled data and question wording affect the results: trained classifiers outperformed Jev on intent classification, a label-trained cascade matched it at 43% of Jev's cost, and a "none" option reduced out-of-scope acceptance from 31% to 19.5%, in independent tests on one benchmark
Labeled data and question wording affect the results: trained classifiers outperformed Jev on intent classification, a label-trained cascade matched it at 43% of Jev's cost, and a "none" option reduced out-of-scope acceptance from 31% to 19.5%, in independent tests on one benchmark

Where AI decision models fit: one stage in the lifecycle of a decision

Rules, trained classifiers, decision models and LLMs can all make a decision in software, each with its own limitations:

Approach

Best suited when

Limitations

Rules and DMN decision tables

The logic can be written down and must be explained rule by rule

Cannot evaluate free text

Trained classifier

The decision is stable, high-volume and has labeled history

Requires labels, and retraining when categories change

Decision model

The options or their definitions change often, and no labels exist yet to train on

No rule to point to; calibration varies by task

Large language model

The output is language, a plan or an explanation

Slower and more costly per decision

Summarized from Kai Waehner's comparison.

These approaches work best as stages in the lifecycle of a decision rather than as competing alternatives. A new decision starts with a decision model, because no labels exist yet to train on. Every answer is logged, and the answers that people confirm or correct become labels.

Once the decision stabilizes, its routine cases move to a less expensive trained classifier. An LLM handles the cases that both the decision model and the classifier are uncertain about, and rules take over the logic that has proven to be fixed.

Decisions that must be explained rule by rule, such as credit or eligibility decisions, remain with a person or a DMN decision table.

A decision moves through stages: a decision model first, a trained classifier once there are labels, an LLM for uncertain cases and rules for what turned out fixed
A decision moves through stages: a decision model first, a trained classifier once there are labels, an LLM for uncertain cases and rules for what turned out fixed

How a decision model checks an AI agent's actions before they run

Inside an AI agent, a decision model can check each action before it runs. Following Kahneman's idea, the LLM plays the role of System Two: it interprets the goal, plans and writes. The decision model plays the role of System One and answers the narrow, repeated questions.

AWS Strands Labs illustrates the pattern with a simple example. Before an agent calls a weather tool with a city the user never mentioned, Strands Decider 2B checks whether the arguments are grounded in the conversation. The agent then asks the user for the missing information instead of guessing.

The decision model checks, policy decides: before a tool call, yes/no questions decide whether the agent proceeds or asks the user, in AWS Strands Labs' example
The decision model checks, policy decides: before a tool call, yes/no questions decide whether the agent proceeds or asks the user, in AWS Strands Labs' example

How to use an AI decision model

This section shows how to serve a decision model, send it questions and act on its answers. The example uses AWS's Strands Decider 2B because it is publicly available for download, while Jev is in early access. Teams install the model and serve it locally:

pip install strands-decider strands-agents strands-decider serve StrandsAgents/strands-decider-2B-hobson-v21 --port 8000

Questions go to the server over HTTP. Each question specifies a type (noul, choice or score) and instructions; choice and score questions also list their options, in a field called criteria. The request format follows Jev's public API, which Kev and Clef also accept, although AWS has not verified compatibility with Jev's API. AWS documents the server for local experiments only: it has no authentication, and AWS does not host it yet.

In an agent built with AWS's Strands Agents SDK, a handler can review each tool call before it runs and return an action such as Proceed, Guide, Deny or Confirm; Confirm pauses the call until a person approves it. An illustrative onboarding gate:

from strands import Agent from strands.interventions import Confirm, Guide, InterventionHandler, Proceed class ActivationGate(InterventionHandler): name = "activation-gate" def before_tool_call(self, event, **kwargs): if event.tool_use["name"] != "activate_account": return Proceed() answers = ask_decider(event) # POST /v1/systemone: two yes/no questions verified = answers["documents_verified"]["noul"] outside_scope = answers["outside_scope"]["noul"] if verified >= 0.9 and outside_scope < 0.5: # illustrative thresholds return Proceed() if verified >= 0.5: return Confirm(prompt="Approve activating this account?") return Guide(feedback="Ask the applicant for the missing documents first.") agent = Agent(tools=[activate_account], interventions=[ActivationGate()])

The decision model judges, the handler decides: activation proceeds when documents are clearly verified, waits for a person when unsure, and goes back for documents when not, with illustrative thresholds
The decision model judges, the handler decides: activation proceeds when documents are clearly verified, waits for a person when unsure, and goes back for documents when not, with illustrative thresholds

ask_decider represents a short HTTP client; AWS's complete example includes one in its repository. The agent's own LLM runs on Amazon Bedrock, which requires AWS credentials with Bedrock access. The thresholds are illustrative: AWS's documentation borrows Jev's example thresholds, acting at 0.9 and confirming from 0.5, and recommends measuring your own. When switching models, re-check the thresholds.

Whichever model an organization adopts, the probability makes it possible to decide in advance which answers the software can act on without human review:

Frequently asked questions

How fast are AI decision models, and what do they cost?

TypeSafe lists Jev at $0.042 per million input tokens, with output free of charge, and 70–500 ms per call measured from the US West Coast. In independent tests, Jev answered in a median of 148 ms at $0.11 per 1,000 requests, compared with 3,609 ms and $1.18 for a general LLM measured in batches on one A100 GPU. Run locally, AWS's Strands Decider 2B answers in about 115 ms on an RTX 3090 GPU, according to AWS. Across a complete system, Kai Waehner expects cost savings of about one order of magnitude.

Are AI decision models simply classifiers?

Not exactly. A decision model receives its options with each request instead of learning them from labels. Once labeled history is available, a small trained classifier can outperform a decision model on intent classification.

Do decision models replace LLMs or AI agents?

No. They handle the narrow, repeated judgments inside an agent or a workflow and leave reasoning, planning and writing to the LLM.

Why not use an LLM's JSON mode instead?

Because acting on an answer requires a confidence level. JSON mode and other structured-output modes constrain what an LLM may write, but the LLM still produces text. A decision model returns a probability per option, designed to be calibrated across many answers. An LLM states a probability in its reply only when asked, and its text must match what the software expects: in independent tests, when an LLM was asked to answer in JSON, 160 of 600 replies did not match the expected format. In either case, an answer chosen from the defined options can still be wrong, and testing on the organization's own data shows how often.

Do decision models work in Spanish?

They need to be validated first. The figures cited in this guide come from English data. Laya's multilingual checkpoint covers more than 100 languages, according to its model card, but its developers found that the English checkpoint could remain confident while failing on some non-English inputs. For other languages, Kev's maintainer, Jared Palmer, recommends a short fine-tune on your own labels.

Can a decision model explain its answer?

No. It returns one value per question and no text, and its logic is learned, so there is no rule to point to.

How is an AI decision model different from DMN?

Decision Model and Notation (DMN) expresses decision logic as tables that a person can read rule by rule. A decision model learns its judgment and returns a probability.

How decision models bring agentic automation to business processes

The question for each organization is where every judgment belongs today: in rules, a trained classifier, a decision model or an LLM. In a business process, an AI decision model and a DMN table can share a single decision step: the model interprets the unstructured input, and the table determines what each answer may trigger. This is how traditional process automation becomes an agentic system, one decision at a time. Aplyca calls this agentic business process management: AI inside business processes, within boundaries the business defines, with people deciding the moments that matter. The process asks, a decision model answers, and, depending on the answer's confidence, the process engine acts on it, passes the case to an LLM or routes it to a person. At Jev's list price, an answer costs about a hundredth of a cent, which makes it affordable to include one at every step of a process. Aplyca helps organizations make this transition and move process orchestration beyond fixed rules alone.

Talk to Aplyca about moving your process automation from rules to agentic systems