AI decision automation: where enterprises can trust an AI model

ARTIFICIAL INTELLIGENCE, BLOG.AUTOMATIZACIÓN.
Abstract 3D render: thin white and lavender planes in a row, crossed by fine violet lines that step to a parallel track at a few planes

In short: AI decision automation lets a model make routine business decisions, such as which team takes a ticket or whether a review is spam, while the platform's rules decide what happens next: the model decides, the platform governs. Dataiku, the data and AI platform company, describes it, also under the name AI decisioning, as a model's judgment combined with business rules and workflow orchestration. AI decision models, a new kind of model, pick one of the business's options and return a probability for each. OpenAI's Decisions API, which serves one of them, is in public beta and charges $0.10 per million input tokens at its base rate. The main caveat: a model given the wrong options still answers, and confidently.

Handing routine decisions to a model can sound ungoverned. This post shows how it stays governed inside the platforms a company already runs, in commerce, content, personalization and business processes, and how to test a model before trusting it. Aplyca delivers marketing technology, personalization, content operations and business process automation (see its case studies), and is a Flowable partner in Colombia and Latin America. Flowable is a platform for business processes and case management. Aplyca's own early test of an AI decision model taught the lesson that runs through this post: the options matter as much as the model.

AI decision automation: traditional rules and AI decision models

AI decision automation lets a model make routine business decisions while the platform's rules decide what happens next. Enterprises already run traditional decision models, which apply rules someone wrote down; an AI decision model judges messy input, such as text or images, and picks one of the options the business defines. The Decision Model and Notation (DMN) standard describes one as "a formal model of an area of decision-making". In practice that usually means DMN decision tables, which an engine applies the same way every time. Each answer can be explained by the rules someone wrote down.

Flowable runs DMN decision tables on typed process variables, for example in a customer onboarding. A DMN decision table's limit is its input: typed values, not text or a photo.

An AI decision model, such as Jev from TypeSafe AI or the model behind OpenAI's Decisions API, reads text, data or images. It returns one of the options the business defines, with a probability for every option. TypeSafe's models also return a confidence, a separate measure of how certain the answer is, which software can set thresholds on.

An AI decision model's answer is a trained judgment, not a written rule: it comes with a probability, not with a rule a reviewer can read. Our guide to AI decision models covers how they work, which ones exist and how they compare with large language models and trained classifiers; this post covers where they fit in the platforms a company runs.

Traditional and AI decision models work together; in the words of Vercel, the web platform company, "Existing decision tables can remain responsible for business policy". The diagram shows that split as we see it.

How we see it: an input such as a review, a visit or a ticket goes to an AI decision model, which picks one of the options the business defines; the platform's rules then decide the action or send it to a person
How we see it: an input such as a review, a visit or a ticket goes to an AI decision model, which picks one of the options the business defines; the platform's rules then decide the action or send it to a person

Each use case below splits the work the same way between the model and the platform:

Use case

The model's question

What the platform keeps

E-commerce

Is this review genuine?

Publish, hide or moderate

Personalization

What is this visitor here to do?

The content an audience sees

Content operations

Which tags apply?

The approval flow

Business processes

Which step comes next?

The DMN decision table, the human task

AI agents

May the agent act?

Approve, confirm or block

In e-commerce, the model sorts listings, reviews and returns, and the store's rules act

A store makes many small judgments: where a listing belongs, whether a review is genuine, why a customer is writing. TypeSafe lists classifying product listings and detecting review abuse among Jev's uses, and its own example sorts products into a standard e-commerce product taxonomy, with no accuracy reported. Kev, a family of open-weight AI decision models by the developer Jared Palmer that a company can run itself, routes a retail customer's message to returns, shipping or billing: AI ticket triage. OpenAI's Decisions API guide checks a product photo for damage.

Laya PHP is an open-source PHP client for Laya, an AI decision model with open weights that a company runs on its own servers. Its review-moderation example shows the pattern: answers with a confidence under 0.7, a score the model returns with its probabilities, go to a person. Laravel, the PHP framework, registers it automatically, and a Symfony or any other PHP application can call it over HTTP.

How we see it: the store's moderation and order rules stay where they are, and the model adds a probability for them to act on.

Aplyca builds online stores and marketplaces, with seller management and "configurable commissions and approval flows", and enterprise applications in PHP with Symfony.

Real-time personalization: one visit, several decisions, and Contentful's rules act

Contentful, the headless content platform, personalizes by audience: a group of visitors "identified by a set of rules". With an AI decision model, the team that runs the audiences keeps them: the model only adds a value to each visitor, and an audience rule of the type "has trait" matches it. Contentful's personalization is deterministic: a visitor who matches the audience always sees the targeted variant. Real time here means the next page view: the model classifies the visitor between page views, and the next page uses the answer.

How we see it: the visitor's recent actions, kept in a cookie, a session store or a customer data platform, go to the model with several narrow questions rather than one intent or lead score. The questions ask what the visitor is here to do, its purpose; whether the visit is commercial; how urgent it is; and how close it is to an action. When an answer is confident, the website sends it to Contentful as that value and the audience rule picks the content; otherwise the visitor keeps the baseline content. The diagram shows that flow.

How we see it: between page views, an AI decision model answers narrow questions about a visitor's recent actions; the answer reaches Contentful as a value that an audience rule matches, and the next page shows a variant, or the baseline when the answer is unsure
How we see it: between page views, an AI decision model answers narrow questions about a visitor's recent actions; the answer reaches Contentful as a value that an audience rule matches, and the next page shows a variant, or the baseline when the answer is unsure

For Aplyca's early test, on synthetic website sessions in TypeSafe's playground, we wrote journeys like these, each with its expected answer written before the model ran. The last column is how we see the site acting on each answer:

The visitor's journey

The answer we expected

What the site does

Services, case studies, pricing, a contact form started

Contact sales; close to acting

Opens the sales path

Support pages, an error page, the status page, a "site down" search

Support; urgent; not commercial

Shows status and support, no sales banner

Careers, culture, team, a job listing

Career interest; not commercial

Shows jobs, no sales call to action

A site search for what an implementation costs

Vendor evaluation

Shows pricing guidance and a case study

A search-engine query for a university assignment

Research; not commercial

No sales follow-up

Pricing three times in a day

Evaluation; rising urgency

Flags the visitor for sales

Incomplete or unreliable session data

Unknown

Keeps the default experience

Contentful's own AI features help people write audiences and variants; none classifies a visitor as a page is served.

In the European Union (EU), the ePrivacy Directive requires consent to read or store information on a visitor's device unless it is strictly necessary. Contentful's software development kit accepts the visitor's consent with each request.

Aplyca implements web personalization and A/B testing for Contentful, with content by "location, behavior and device"; personalization in a digital experience platform covers the basics.

In content operations, the model tags and screens, and the editorial workflow approves

AI decision models can tag a content entry and flag whether it needs review before it is published. TypeSafe documents moderation as combining "severity and confidence to allow, warn, review, or block content". In its own test classifying 60 company annual reports into industry groups, 27 of the 30 answers given with a confidence of 0.9 or more were right; below that, 12 of 30 were. Databricks, the data platform, offers a similar function in beta,

, with examples such as tagging customer reviews and asking "Does this document need human review?"

How we see it: the model proposes tags and flags risk, and the content management system (CMS) or digital experience platform (DXP) keeps the decision: its workflows decide who signs off, and a low-confidence answer adds a reviewer.

Aplyca implements headless content platforms such as Contentful, with "versioning and workflows" and "multi-level approval flows".

AI document classification in Flowable: the model answers, DMN applies the policy

In a business process, someone has to act on the model's answer, within rules, with a record. Flowable does that job with processes in Business Process Model and Notation (BPMN), cases in Case Management Model and Notation (CMMN) and decisions in DMN. Its Agent Engine, the runtime for AI agents, works beside them.

Take customer onboarding, where each incoming document needs AI document classification before the case can move; here, classifying a document means choosing the step it goes to. Kai Waehner, an enterprise architect, describes a confidence gate in a workflow, with DMN as the alternative for decisions that must be explained by rule. How we see it, DMN is where the policy lives:

  1. The case asks, the model answers. The case sends the document to the AI decision model. Its options are the plan items open at that point, the steps the case may take next, plus "needs review". The case model, in CMMN, is the harness that limits what both the AI decision model and the Orchestrator Agent, Flowable's coordinating agent, may choose.

  2. Confidence decides who decides. Against the thresholds the DMN decision table holds, a confident answer files the document and moves the case on. A middling one goes to the Orchestrator Agent, and a doubtful one to a person.

  3. Probabilities become variables, and DMN applies the policy. A DMN decision task, the case step that runs a DMN decision table, reads the probabilities and the confidence as typed case variables and holds the thresholds.

  4. Cheap checks run on every step. Flowable 2026.1 adds guardrails, a safety layer around an agent's input and output, and agent evaluators, which score its result. How we see it, an AI decision model called as a service is one more check of that kind.

  5. The audit trail becomes the training set. Each decision, with the person's correction where there was one, becomes a labeled example. Those examples evaluate the model and, once there are enough, train a cheaper classifier for the routine documents, the next stage in a decision's lifecycle, as our guide describes it; the audit trail also feeds process mining.

The diagram shows the onboarding flow as we see it.

How we see it: in a Flowable onboarding case, an AI decision model classifies an incoming document into one of the steps open now; a DMN decision table applies the business's thresholds and sends the case on, to the Orchestrator Agent or to a person
How we see it: in a Flowable onboarding case, an AI decision model classifies an incoming document into one of the steps open now; a DMN decision table applies the business's thresholds and sends the case on, to the Orchestrator Agent or to a person

Flowable has no connector built for AI decision models, but a case can call one as an external service, hosted or run on the company's own servers: an HTTP task or a REST operation from Flowable's service registry sends the document to the model's API and keeps the answer as case variables that the DMN decision task reads. A case's decision task can also be "AI Activated" so that the Orchestrator Agent can use it. How we see it: the AI decision model's answer gives the Orchestrator Agent a fast, inexpensive first judgment before it reasons at length; our guide to AI decision models calls that role a System One. The case keeps a fallback for when the model does not answer.

AI ticket triage follows the same pattern: the AI decision model picks the team for each support message, and the platform's rules decide whether to route it on that answer or send it to a person. AWS's Strands Decider, an AI decision model, routes a support message about failing payouts to billing in its documentation.

As a Flowable partner in Colombia and Latin America, Aplyca implements Flowable for processes and cases with approvals, audits and AI agents, and automates sales, onboarding and support workflows with n8n, a workflow automation tool.

Before an agent acts, a cheap check decides whether it may

AI agents that act on their own need the same split between the model's judgment and the platform's rules. TypeSafe, AWS with its Strands Decider model, and Databricks each document a check that runs before an agent acts or after it answers. Our guide walks through AWS's check in code, and our post on the Model Context Protocol covers how agents reach enterprise systems. How we see it: the platform asks the model an inexpensive question before each action, and decides what a "no" does.

Aplyca builds generative AI solutions for businesses, from assistants "for service, sales and support" to AI agents.

When the right option is missing, the model still answers

An AI decision model always picks from the options it is given. Aplyca's early test on synthetic website sessions showed what that means: a visitor reading careers pages was classified as a researcher, confidently, because the options had no careers purpose. The redesign added a careers purpose and asked one question per concept, such as purpose, urgency and closeness to an action. It also separated "other", for a purpose that fits none of the options, from "unknown", for too little evidence.

The vendors give the same advice. OpenAI recommends a fallback such as "other" that goes to a review queue. TypeSafe recommends an "other" or "none of the above" option with a description for each. Our guide covers how the wording of a question changes the answers.

The options deserve the care of a policy, because in practice they are one.

Count the cost of a governed decision, not of a model call

The cost that matters is the cost of a governed decision. How we see it, the total is the price of the call plus the cost of review: the share of answers the threshold sends to a person, times what each review costs. In TypeSafe's annual-report test, a threshold of 0.9 would have sent 30 of 60 answers, half, to a person: at that share, the cost of review, not the price of the model, decides the total.

OpenAI's Decisions API charges a base rate of $0.10 per million input tokens, the units of text it bills by, and nothing for output; premiums apply for regional processing and long inputs. Laya, run through Laya PHP, has no per-call charge; the cost is the server.

HatchWorks, a consultancy, checks whether each answer is valid, correct and authorized, and each "no" adds a review or an error. Dataiku tracks error rate, cost per decision and how often people override the model.

Thresholds, model versions, vendors and consent are policy

The business that owns the risk sets the thresholds, the model version, the vendor and the consent rules. Dataiku recommends shadow mode: running the model beside the current process first, with a fallback to a person. Thresholds belong with the workflow, versioned with it and tied to one model version; our guide covers setting them on labeled data.

Inputs are policy too: send the model only the fields a question needs. So is the vendor: Jev is a hosted service with no published service-level agreement, while Clef, Cloudflare's AI decision model, and Kev publish open weights a company can run itself.

In the EU, classifying visitors from their behavior is profiling under the General Data Protection Regulation (GDPR), and visitors may object to it. The ePrivacy Directive's consent rule also covers the cookies that hold their actions. Which rules apply to a given use is a question for counsel; in Colombia, that includes Ley 1581 de 2012.

Test an AI decision model before you trust it, and skip it where a rule already answers

An AI decision model is the wrong tool when the answer must be free text rather than one of a set of options, when the options cannot be defined in advance, or when a rule already captures the decision.

Where it fits, start with one decision point. Write the expected results before running the model, change one variable at a time, and run the model in shadow mode beside the current decisions. Aplyca's test was redesigned to change one variable at a time after its first test cases proved too loose to measure. In the redesign, the same journey runs slow, then fast, then fast with a contact form started, and urgency should rise in that order. Trap cases check that a pricing page alone, a contact page opened for the office address or a burst of fast clicks does not read as buying intent.

A perspective for Latin America: test in Spanish and know where the data goes

No published evaluation measures AI decision models on Spanish-language decisions, so test any model on your own Spanish or Portuguese test cases first. Our guide sums up what the models' makers say about other languages.

OpenAI processes Decisions API data in the United States or Europe and names no Latin American region.

In Colombia, the personal data law, Ley 1581 de 2012, requires prior, express and informed authorization to process personal data and restricts transfers to countries without adequate protection. The Superintendence of Industry and Commerce, the country's data protection authority, treats data collected through cookies as personal data.

Frequently asked questions

Do we need new platforms to use AI decision models?

No. The model is called as a service, and each platform keeps its rules. A PHP store calls it through Laya PHP or over HTTP, a website passes its answer to Contentful as a value that an audience rule matches, and a Flowable case calls it as an external service.

Can Flowable call an AI decision model?

Yes, by configuration rather than a built-in connector. A case sends the document to the model's API with an HTTP task or a REST operation from Flowable's service registry and keeps the answer as case variables, and a DMN decision task applies the thresholds. For the language models its AI agents use, Flowable accepts models from OpenAI, Anthropic or Azure OpenAI, or any model behind an OpenAI- or Anthropic-compatible API, including one a company hosts itself.

Can we run an AI decision model on our own servers?

Yes, with an open-weight model. Laya runs on a company's own servers with no per-call charge, and Kev and Clef publish open weights. Jev and OpenAI's Decisions API are hosted services. Where personal data has to stay in the country, an open-weight model is the one to test first.

Can AI decision models be used in regulated processes?

Yes, when the process keeps the rules and the record. The EU AI Act asks high-risk systems to log events automatically and to allow human oversight that can override or stop them. A process that logs each decision and sends doubtful ones to a person supports that oversight. In Colombia, Ley 1581 de 2012 also governs the personal data such a process uses; elsewhere, including the United States, the rules are a question for counsel.

Can an AI decision model personalize a page while it loads?

Not on published evidence. TypeSafe states about 100 ms for most Jev queries without saying under what conditions, which is not a measurement of a page load. Classifying between page views avoids the wait.

AI decision automation starts with one routine decision and the rules to keep

For each team, the question now is which routine decision to give a model and which rules the platform keeps. The model decides, the platform governs: the options, the thresholds, the reviewers and the record stay with the business. The best first candidate is a frequent decision that people make today from messy input, inside rules that already decide what happens next, such as the visitor's purpose on a website that personalizes with Contentful, or the step an incoming document takes in a Flowable onboarding case.

See how Aplyca sets up personalization and A/B testing on Contentful

See how Aplyca implements Flowable for processes and cases