Short answer
A model can play three roles in a decision: describe what happened, predict what is likely, or recommend an action. Only the third changes anything, and only a narrow class of decisions should be handed over entirely — explicit rule, measurable outcome, cheap to reverse. Everything else gets prepared, not decided. The constraint is almost never the model: it is whether your data is consistent enough to match a customer across two systems. And the only honest measure of success is a decision that got made differently, not a report that got delivered faster.
The three roles a model can take in a decision
Conversations about AI and decision making get muddled because three very different things share the name. Separating them makes the rest of the question tractable.
| Role | What it produces | Who decides |
|---|---|---|
| Description | What happened, in language rather than a chart. Cheap, useful, least likely to change behavior — the territory of automated reporting, and it stops there. | A person, unaided. |
| Prediction | What is likely next, given history. Valuable where the business is stable enough for history to mean something, close to worthless where it is not. | A person, better informed. |
| Recommendation | What to do about it, with the reasoning attached. | A person — unless the decision passes all three conditions below, in which case the model can. |
Most projects sold as decision support deliver the first, hint at the second, and never reach the third. That is why the dashboard goes unopened after six weeks: nothing about it required anyone to do something differently.
The decisions you can genuinely delegate
A decision is a candidate for full delegation when three conditions hold at once. Each is checkable before you build anything.
- The rule is explicit. You can write down what should happen in each case, even if the number of cases is large. If your best experts disagree about the rule, a model will not resolve the disagreement — it will pick a side silently.
- The outcome is measurable soon. Within days, ideally. A decision whose quality only becomes visible in a year cannot be improved by a feedback loop, because there is no loop.
- A wrong call is cheap to reverse. Reordering stock, flagging a document, routing an enquiry. Not signing, not terminating, not pricing a contract.
Applied honestly, that filter leaves a smaller list than most companies expect — and a much more reliable one. A related but distinct filter, four criteria rather than three, decides which process to automate first; it is worked through in our process audit method.
The ones that must stay with a person
Some decisions fail the filter for reasons that no amount of model quality changes.
- Anything involving a specific person — hiring, promotion, discipline, individual credit terms. Beyond the legal exposure, these are the decisions where an unexplainable output is least acceptable to the person on the receiving end.
- Commitments with a long horizon. Choosing a market, a supplier you will depend on, a product direction. The feedback loop is too slow to correct an error.
- Anything you would have to justify to a client, an auditor or a regulator. The standard there is not accuracy, it is being able to explain the reasoning, and a recommendation you cannot reconstruct fails that test even when it is right.
- Decisions where the cost of being wrong is asymmetric. If a false positive costs an email and a false negative costs a client, the model needs a human on the expensive side of the asymmetry.
The useful reframe: a model is excellent at making sure a decision is well prepared and terrible at being accountable for it. Design around that split and most of the governance questions answer themselves.
What a forecast is worth at this size
Forecasting is where expectations diverge most from reality. The determining factor is not the technique — it is how much of your revenue is repeatable.
A company with subscription or recurring revenue and recognizable seasonality can get a useful forecast from a modest history, because the past genuinely constrains the future. A company whose quarter depends on whether two large deals close will not get a reliable forecast from any method, and a model that produces one anyway is producing false precision. In that second case the right output is not a number but a set of scenarios with the assumption behind each made explicit.
The bar to clear is comparative, not absolute: a forecast earns its place when it beats the estimate you were already making informally. Track that comparison for a couple of months before deciding whether it is working — the specific case of cash is covered in our guide to cash flow forecasting.
Telling a good recommendation from a confident one
Language models produce fluent, assured prose regardless of how thin the evidence underneath is. That is the single most dangerous property in a decision-support context, because fluency is exactly the signal humans use to judge competence.
Three practices from LYVIA’s own engagements make the difference visible, offered as working habits rather than published benchmarks:
- Require the inputs, not just the answer. A recommendation should name the figures it used. If it cannot, it is a rewording of the prompt.
- Ask for the case against. Prompting for the strongest argument on the other side surfaces thin reasoning faster than any confidence score — a technique described further in our guide to prompt engineering for business.
- Check it against decisions you already made. Run the system over last quarter, where you know the outcome. Agreement with the good calls and disagreement with the bad ones is the only validation that means anything.
The plumbing that decides everything
Almost every stalled decision-support project stalls in the same place, and it is not the model. The information needed to make the call lives in three systems that cannot agree on what a customer is.
What unblocks it is unglamorous and worth doing regardless: one system of record per kind of information, a shared identifier so records can be matched, and access from the automation layer without a manual export. Companies that fix this once find that the second and third use cases cost a fraction of the first. Those that skip it end up with a model reasoning confidently over data that contradicts itself — the architecture side of this is covered in our guide to AI infrastructure.
The decision log, and why it is the whole system
One artifact separates the companies where this works from the ones where it quietly does not: a record, per decision, of what was recommended, on what inputs, and what the person actually did.
It costs almost nothing to write and it is the only thing that answers the question you will care about in six months — is this helping? Without it you have anecdotes. With it you can see that the system is overridden nine times out of ten in one category and never in another, which tells you precisely where to narrow the scope and where to widen it. It is also what makes the decision defensible to a client or an auditor, because the reasoning was recorded at the time rather than reconstructed afterwards.
Where to start
Pick one recurring decision that already gets made weekly, where you can name the person who makes it and the information they wish they had. Run the model alongside them without authority for a few weeks, logging both. Compare. Then either give it a narrow mandate inside that decision or drop it and pick another — both outcomes are cheap at that point, which is the entire reason to sequence it this way.
What does not work is starting with the decision that matters most. The stakes make people override on principle, you learn nothing, and the project acquires a reputation before it has had a chance. Where this sits against the rest of a first year is set out in our AI strategy roadmap.
Frequently asked questions
Can AI make business decisions for us?
It can make some of them, and the useful question is which. A decision qualifies when the rule is explicit, the outcome is measurable within days, and a wrong call is cheap to reverse — reordering stock below a threshold, flagging an invoice for review, routing a lead. Decisions that involve people, commitments over a long horizon, or anything you would have to defend to a client or a regulator stay with a person. The value of a model in that second category is preparation, not authority.
What is the difference between a dashboard and AI-supported decision making?
A dashboard tells you what happened. It leaves the interpretation entirely to whoever opens it, which is why so many go unread. AI-supported decision making adds two things a dashboard cannot: an explanation of why a number moved, and an explicit recommendation with its reasoning attached. The recommendation is what makes it useful and also what makes it dangerous — a confident recommendation built on a weak signal reads exactly like a strong one.
How accurate are AI forecasts for a small business?
Accurate enough to change a decision, rarely accurate enough to be trusted unexamined. Forecast quality depends far more on how stable your business is than on the model: a company with recurring revenue and steady seasonality gets useful numbers from very little history, while one whose revenue depends on a handful of large deals will not get a reliable forecast from any technique. Judge a forecast by whether it beats your current guess, not by whether it is precise.
What data do we need before AI can support decisions?
Less than most vendors imply, but it has to be consistent. One system of record per kind of information, a shared identifier so a customer in one system can be matched to the same customer in another, and enough history to cover at least one full cycle of whatever you are deciding about. Scattered data is the usual blocker, not volume — and it is a plumbing problem, not a modelling one.
Who should be accountable when a decision follows an AI recommendation?
The person who would have been accountable without it. Nothing about a model transfers responsibility, and treating a recommendation as cover is the failure mode to watch for. What changes is what you record: the recommendation, the inputs behind it, and whether the person followed or overrode it. That log is what lets you find out months later whether the system is actually helping.
Where do companies most often go wrong with this?
They automate the reporting and call it decision support. The numbers arrive faster, nobody acts differently, and the project is judged a success because the dashboard is live. A decision-support system that does not change a single decision has failed, however clean the pipeline is — which is why the metric to agree on up front is a decision that got made differently, not a report that got delivered on time.
If you want to know which of your recurring decisions could genuinely be prepared or delegated — and which should not be — that is the conversation to have first. Book a call.
