Reference Guide · AI

AI in Production Systems — Engineering, Not Demos

Building an AI demo is easy; running an AI system reliably is engineering. Between the two lies everything that matters: evaluation, failure modes, boundaries, cost. How to put AI into production instead of putting it on show. A decision document for CTOs, founders and senior engineers.

What is this? · Reference Guide

A solid guide to an engineering question — with trade-offs, costs and the case in which we decide differently. Not an opinion piece, but a reference text. Go to overview

Author
Batunet Engineering
Reading time
16 min
Level
In depth
Status
Approved
Last reviewed
21 July 2026
Updated
21 July 2026
On this page

There is hardly a field in which the gap between first impression and reliable operation is as wide as it is with AI. A demo comes together in hours and impresses immediately; a system that delivers the same capability in production reliably, affordably and without silent errors is one of the more demanding engineering tasks. The market shows demos; what comes after them is rarely discussed honestly.

This document is about what comes after. It treats AI neither as magic nor as a threat, but as a special kind of component — one that gives probable rather than certain answers — and asks how to build a system around it that you can take responsibility for. It is deliberately kept neutral with regard to models, vendors and hype; it names no product, because the principles depend on none.

1. The gap between demo and production

A demo has a single job: to show that something works in the best case. It is presented with hand-picked examples, under friendly conditions, without load, without edge cases, without the question of what happens when the answer is wrong. A production system has the opposite job: to be reliable, especially in the bad case, under real load, with real users whose inputs nobody anticipated.

Demo Evaluation · guardrails · context Cost · latency · error handling the actual system visible all the work

Diagram: the demo is the tip above the waterline. Everything that makes an AI system reliable lies beneath it — invisible at first glance.

The gap between the two is especially deep with AI, because the demo succeeds so easily and says so little about the production system. A model that convinces in ten showcased cases can silently assert something false in the eleventh — and that eleventh case is the one that matters in production. Anyone who takes the demo's success as proof of production readiness confuses the best case with the normal case. The real work begins exactly where the demo ends.

Trade-off. Taking the demo seriously as what it is — a beginning, not a proof — dampens the enthusiasm it sparks and requires naming the long road behind it.

Cost. The distance between demo and production is work that is invisible at first glance and therefore easily forgotten in the budget.

When we decide differently. Where an AI capability really is only meant to be shown — an experiment, a learning exercise — the demo is enough, and the effort of production readiness would be misplaced.

2. Why AI systems are different

Conventional software is deterministic: the same input leads to the same output, and if it doesn't, that is a bug you find and fix. An AI model is probabilistic: it gives the most plausible answer, not the certain one, and the same question can produce different answers. That is not a bug but the nature of the thing — and it changes everything you know about working with the component.

The rest follows from this one property. You can't test a probabilistic system like a deterministic one, because there is no fixed correct output to check against. You can't rely on it doing the same thing twice. You have to expect it to be convincingly wrong — not with an error message, but with a fluent, plausible, wrong answer. Anyone who treats an AI system like ordinary software applies tools built for certainty to something that delivers probability — and is surprised when they don't work.

PropertyDeterministic softwareAI component
Outputfixed, reproducibleprobabilistic, variable
Errorsfails visiblyoften silently and plausibly wrong
Verificationtest the expected outputevaluate across many cases
Reliabilityinherent at the corehas to be actively established

Trade-off. Accepting probability means giving up the convenience of the deterministic — you trade certainty for capability and have to establish the certainty elsewhere.

Cost. Working with a probabilistic system requires methods that many teams first have to learn — evaluation, safeguards, monitoring for the unexpected.

When we decide differently. Where a task can be solved with certainty and by rules, you don't use a probabilistic system — deterministic software is simpler, cheaper and more reliable there (see also the question of when AI is the wrong solution).

3. Evaluation, not impressions

Because an AI system has no fixed correct output, you can't judge its quality by a single case — least of all by the showcased one. You need evaluation: a set of cases with known, desired results against which you systematically measure the system, and a measure of how often and how far it deviates. Without that measurement, every statement about the quality of an AI system is a feeling, not knowledge.

Evaluation is to AI what testing is to ordinary software — only harder, because the correct answer is rarely unambiguous. It is therefore not a one-off acceptance check but a permanent fixture: before every change to the system, you measure whether it has gotten better or worse, because with a probabilistic system a well-intentioned adjustment in one place can degrade something in another without anyone noticing. Anyone who tinkers with an AI system without evaluation is working blind — changing and hoping instead of changing and knowing.

Trade-off. Building evaluation takes time and care before the system delivers any value at all — it feels like a detour, but it is the foundation of every reliable statement.

Cost. An evaluation set has to be maintained, grow with the cases and cover the edge cases where the system fails — that is ongoing work, not a one-time effort.

When we decide differently. For a throwaway experiment whose errors have no consequences, you can skip full evaluation; but as soon as the system influences decisions that matter, it is indispensable.

4. Failure modes of AI systems

An AI system fails differently from ordinary software, and its failure modes are dangerous precisely because they don't look like failures. The best known is the convincing false statement: the model asserts something plausible that isn't true, in fluent, self-assured form — without any warning that it is guessing. Such an error doesn't abort anything; it flows into the answer and is believed, because it is indistinguishable from a correct answer.

Then there are quieter forms. Behavior can drift over time as inputs or the model shift, so a system that was good yesterday is worse today without anyone having changed anything. It can fail on rare but important inputs that never appeared in any demo. And it can be pushed out of its intended role by cleverly crafted inputs. These failure modes have one thing in common: they don't report themselves. You only find them if you actively look for them — through evaluation, through monitoring, through deliberately designing for the case where the answer is wrong.

Failure modeWhat happensCountermeasure
Convincing false statementplausible but wrong, without warningevaluation, checking against sources, human oversight
Driftquality declines unnoticed over timecontinuous measurement, re-evaluation
Edge-case failurerare inputs break qualityadd edge cases to the evaluation
Input manipulationsystem leaves its roleboundaries, guardrails, don't trust inputs blindly

Trade-off. Designing for the failure modes means distrusting the very system you are deploying — this sobriety dampens the enthusiasm, but it is the prerequisite for reliability.

Cost. Every countermeasure — checks, monitoring, oversight — is additional work and additional parts in the system that themselves have to be maintained.

When we decide differently. Where a wrong answer has no consequences — a suggestion that a human reviews anyway, a non-critical extra — you can slim down the elaborate countermeasures; full safeguarding applies where the answer triggers something.

5. Guardrails and boundaries

A probabilistic system needs boundaries within which it may act — and outside of which it cannot act. These boundaries are themselves deterministic: fixed rules that check what goes into the model and what comes out, and that prevent a wrong or unauthorized answer from taking effect. The model proposes; the boundary decides whether the proposal may be carried out. That way, responsibility stays with something verifiable, not with something probabilistic.

deterministic enclosure Input validated Model (probabilistic) Output validated

Diagram: the probabilistic core is enclosed by deterministic boundaries. What goes in and comes out is checked — responsibility stays with something certain.

Where the impact demands it, the human is part of the boundaries too. Not every AI output has to be checked by a human — that would destroy the benefit — but wherever a wrong answer causes serious harm, human oversight belongs in the path between proposal and effect. The art is to apply this oversight where it matters and leave it out where the error has no consequences — not everywhere and not nowhere.

Trade-off. Boundaries and oversight buy reliability with part of the speed and autonomy that AI promises — a completely unconstrained system is faster, but irresponsible.

Cost. Every boundary is additional deterministic logic that has to be designed, tested and maintained, and every human check costs time in the workflow.

When we decide differently. Where the model's output is only a non-binding suggestion to a human anyway, you can do without additional boundaries — the human is then already the boundary.

6. Data and context are the system

The quality of an AI system in production depends less on the model than on what you give it. A model is a general-purpose tool; it only becomes useful through the context you feed it — the right, current, reliable information for the right question. An excellent model with poor context gives poor answers; a modest model with good context often gives surprisingly good ones. That is why the real engineering work often lies not with the model but with the context: sourcing it, keeping it current, getting it to the model cleanly and relevantly.

This shifts the focus from the question "which model?" to the more important "which context, and from where?". Is the information source reliable and current? Is the right slice selected for the specific question, not too much and not too little? Does it remain traceable what an answer is based on? These questions determine quality in production — and they are classic data engineering, not magic. Whoever takes them seriously holds the biggest lever for quality.

Trade-off. Investing in context rather than in the model means doing the less glamorous work — data maintenance instead of model selection — but pulling the most effective lever.

Cost. Providing good context and keeping it current is ongoing work on data sources, selection and maintenance — it never ends, because the world changes.

When we decide differently. Where the task needs no external knowledge but relies purely on the model's capability, context takes a back seat and the choice of model carries more weight.

7. Cost and latency as an architecture question

Unlike ordinary software, whose per-request operating cost is often negligible, every AI request has a noticeable price — in compute time, in money, in waiting time for the user. These costs are not a side issue to optimize later but an architecture question from the start. A system that queries the most expensive model for every little thing can impress at small scale and become unaffordable at large scale.

The responsible approach treats the model like an expensive, slow resource to be used sparingly. You query it only where it is really needed, not everywhere; you hold on to answers where the same question recurs; you choose the smaller, cheaper option for simple tasks and the stronger one for hard ones; you plan for what happens when the model is slow or fails. Cost and latency are thus not operational details but shape the architecture as strongly as the database shapes an ordinary system.

Trade-off. Taking cost and latency seriously means not using the strongest model everywhere — you trade a little capability for affordability and speed.

Cost. The economical architecture — tiered models, caches, fallback paths — is itself additional complexity that you build and maintain.

When we decide differently. Where requests are rare and the answer is especially valuable, you may choose the most expensive option; the economy applies where volume and repetition drive the cost.

8. Determinism where it counts

The idea connecting all of these chapters is an architectural one: you wrap the probabilistic in the deterministic. The probabilistic part — the model — is confined to the area where its capability is the only way forward; everything around it that can be certain is built to be certain. Checks, boundaries, context selection, output handling: all of this is ordinary, deterministic software that you can test and take responsibility for. The system as a whole thereby becomes controllable, even though its core is not.

That resolves the apparent contradiction of using an uncertain tool responsibly. You don't make the model safe — you can't — but the system around it. The probabilistic core remains what it is, but it is enclosed, monitored and bounded, so its uncertainty stays local and doesn't permeate the whole. This is the same attitude that applies everywhere in good engineering: keep the uncertain small and put it behind something verifiable.

Trade-off. Enclosing the probabilistic costs the additional deterministic structure around it — you build more than just the model call in order to be able to take responsibility for it.

Cost. The enclosure is real work: checks, boundaries and monitoring are components in their own right with their own maintenance, multiplying the effort well beyond the bare model call.

When we decide differently. Where the model's uncertainty has no consequences, you can skip the full enclosure; the care grows with the harm a wrong answer can cause.

9. Common mistakes

The recurring patterns that make AI fail in production — almost all of them are variants of mistaking the demo for the system:

  • Taking the demo's success as proof of production readiness and underestimating the road behind it.
  • Treating a probabilistic system like deterministic software and relying on fixed, reproducible outputs.
  • Tinkering with an AI system without evaluation — changing and hoping instead of changing and measuring.
  • Overlooking the silent failure modes because they don't look like failures but like fluent answers.
  • Trusting the model blindly instead of wrapping its inputs and outputs in deterministic boundaries.
  • Optimizing the model where the context is the real problem — trying to replace good data work with model selection.
  • Treating cost and latency as a later optimization until the system becomes unaffordable or too slow at scale.
  • Applying human oversight everywhere or nowhere instead of where a wrong answer really does harm.

10. Decision checklist

Before and during the build of an AI production system, clarify the following in order:

  • Demo or system? Is it clear that the impressive demo is only the beginning — and has the road to reliability been planned for?
  • Probabilistic nature understood? Is everyone aware that outputs are variable and sometimes convincingly wrong — and is the system designed for that?
  • Evaluation in place? Is there a maintained set of cases against which quality is measured — before every change?
  • Failure modes considered? Are convincing false statements, drift, edge cases and input manipulation addressed?
  • Enclosed? Is the probabilistic core surrounded by deterministic boundaries that validate input and output?
  • Human where needed? Is there human oversight in the path where a wrong answer causes serious harm — and only there?
  • Context under control? Is the information source reliable, current and traceable — and is the right slice selected?
  • Cost and latency planned? Has the price per request been considered, with tiered models, caching and fallback paths?
  • Observable? Can you see in production how the system behaves — including silent degradation?

If you can answer these questions, you have built an AI system you can take responsibility for — not just a demo that impresses.

FAQ

Why isn't a convincing demo enough? Because a demo shows the best case, and a production system has to survive the bad one. A model that shines in ten showcased cases can silently assert something false in the eleventh — and that eleventh case is the one that matters in production. The real work begins where the demo ends.

How do you test a system that doesn't always give the same answer? Not with one expected result, but with evaluation across many cases: a set of examples with desired results and a measure of how often and how far the system deviates. You measure quality statistically, not by the individual case, and you measure again before every change, because an improvement in one place can do harm elsewhere.

What is the most dangerous failure mode? The convincing false statement. The model delivers a fluent, plausible, self-assured answer that isn't true — without warning, without aborting. It is dangerous because it is indistinguishable from a correct answer and is therefore believed. Countermeasures are evaluation, checking against reliable sources and human oversight where it matters.

Does the model or the data matter more? Usually the context you give the model. A strong model with poor context answers poorly; a modest one with good context often answers remarkably well. The most effective work therefore rarely lies in model selection but in getting reliable, current information to the model cleanly and relevantly.

Does a human have to check every AI output? No — that would destroy the benefit. Human oversight belongs where a wrong answer causes serious harm and can be dropped where the error has no consequences. The art is to apply it selectively, not everywhere and not nowhere.

How do you keep costs under control? By treating the model like an expensive, slow resource: query it only where it is needed; hold on to recurring answers; choose the smaller option for simple tasks; plan a fallback for when it is slow or fails. Cost and latency are an architecture question from the start, not a later optimization.

Further reading

It is grounded in the Batunet Engineering Method: enclose the uncertain, design for failure, observability from day one, keep judgment with humans.

Closing engineering principle

AI doesn't change what makes good engineering — it only puts it to the test. A probabilistic model is no reason to abandon discipline but the strongest reason to apply it: measure instead of believe, design for failure, put the uncertain behind something certain, leave judgment with humans. The difference between an AI demo and an AI system is exactly the difference between impression and engineering. The model provides the capability; the reliability comes from the system you build around it. Whoever understands this sees AI not as a shortcut around the craft of engineering, but as a new, demanding application of it.


A demo shows that something can work. A system makes sure it works reliably — even in the bad case. Between the two lies all the work, and it is engineering, not magic.

Referenced entities

Knowledge graph

Continue your engineering journey.

Related concepts, decisions, playbooks and perspectives — as one connected path, not a list of links.

A concrete project in this field?

Reference Guides show how we think. For your system, talk to our management — technical, no sales pitch.