Reference Guide · AI

Agentic Coding — When Software Development Becomes Orchestration

An agent that reads a repository, plans in multiple steps, changes many files, runs tests and reacts to their results doesn't take engineering work off your hands — it shifts it. Where that work moves, what becomes more expensive there and what boundaries a system needs when work is done this way. A decision document for CTOs, engineering leads and senior engineers.

What is this? · Reference Guide

A solid guide to an engineering question — with trade-offs, costs and the case in which we decide differently. Not an opinion piece, but a reference text. Go to overview

Author
Batunet Engineering
Reading time
20 min
Level
In depth
Status
Approved
Last reviewed
4 September 2026
Updated
4 September 2026
On this page

The conspicuous change is easy to name: where people work with AI agents, they report typing less themselves. It is also the one people talk about. The consequential one shows up later and more quietly — it concerns not how code comes into being, but where work on a system gets stuck.

The question is therefore not what an agent can do. It is: when generating code becomes cheaper — where does the work get stuck then? The answer is not “nowhere”. Our thesis — and it is ours, not a finding — is this: the bottleneck shifts away from producing code toward defining, orchestrating, verifying and taking responsibility for work. Exactly where it moves is described in the next chapter, along a division that is likewise ours. The text names no product, no model and no figure. What matters here doesn't change with the state of the tools — and we know of no reliable figure for this way of working that we would stand behind.

1. What actually shifts

What used to happen on the side — deciding what should be built, and being certain that what was built is right — becomes the actual work. The perceived gain in generation is therefore not the whole calculation: the effort doesn't get smaller, it is distributed differently. The objection that defining and verifying were always part of the job is correct. What is new is not the activity but its share. Whether a by-product also falls away in the process — whether knowledge forms differently in building than in verifying — is an open question; chapter eight comes back to it.

We describe work on a system along four stations: defining, producing, verifying, taking responsibility. This division is our way of thinking about it, not a research result. As long as producing was the most expensive station, the work got stuck there, and everything else arranged itself around that bottleneck. If producing becomes cheaper, the constriction doesn't dissolve — it changes place.

until now when generating becomes cheaper the familiar bottleneck the new bottleneck the new bottleneck the new bottleneck Defining Producing Verifying Responsibility

Diagram: our way of thinking about it — the bottleneck doesn't disappear, it changes place.

Trade-off. Accepting the shift means offsetting the perceived gain in generation against an invisible increase in defining and verifying — the calculation turns out more sober than the first impression.

Cost. The new work is poorly visible and therefore rarely planned for; it appears in no estimate that assumes the old shape of the work.

When we decide differently. Where a task is small, isolated and without consequences, the shift doesn't weigh much — the new work hardly arises there, and you get what the way of working promises anyway.

2. The difference from autocompletion and chat

A completion suggests, a chat answers; in both cases the human accepts or rejects and executes. An agent acts: it reads in the repository, breaks a goal down into steps, makes changes across several files, calls tools, reads their results and continues with them. This feedback from the environment is the distinguishing feature: the next step depends on what the previous one achieved. We distinguish the three forms along five characteristics — the distinction is ours, not a classification from research:

DimensionCompletionChatAgent
Contextthe spot where you're writingwhat is put inrepository, tools, execution results
Scope of one actionone suggestionone sectionsequence of changes across several files
Feedbacknonethe human repliesresult of execution and tests
What has to be verifiedone line in viewone adopted sectiona sequence of changes including state
Who executeshumanhumansystem, within the granted permissions

Our reading of the difference: the leap lies in the scope of action, not in the quality of the model behind it. Whoever objects that this is just a better assistant overlooks who executes the intermediate steps: as long as a human triggers every step, they are also its reviewer. Once that falls away, verifying becomes a work step of its own.

Trade-off. As the scope of action grows, so does what a tool can take over — and, in the same move, the amount of what nobody has read along.

Cost. A sequence of changes across several files can't be skimmed like a suggestion — it has to be traced through, and that costs scarce attention.

When we decide differently. Where a task consists of a single, manageable step, the smaller scope of action is the more fitting choice — not the weaker one.

3. Defining work: tasks, acceptance criteria and boundaries

If you can't say how “done” can be recognized, you can't delegate a task. Breaking down tasks and defining acceptance criteria are thus no longer preparatory work but the accomplishment itself. The objection that this is waterfall through the back door misses the point: what gets fixed is not the solution but the acceptance. The path stays open; the goal has to be verifiable. With parallel work, this applies all the more: parallelism multiplies the burden of defining and verifying instead of eliminating it — the bottleneck is reached sooner, not later. This cutting to size of what is to run side by side is the orchestration the title speaks of.

The second place where things get fixed is the system itself. We consider it decisive that boundaries sit where they take effect mechanically: in module boundaries, types, contracts, tests and conventions — not in the task description. A boundary in the description is cheap and doesn't hold; one in the structure is expensive and bears weight. It does two things: it is the feedback by which an agent recognizes whether a step held, and the gate at which a human verifies. A repository whose rules can be checked mechanically guides; one whose rules live implicitly in a few people's heads leaves both to chance.

Trade-off. Definition and structure cost time precisely where the way of working promises to save time — the perceived gain shrinks as a result.

Cost. Structure is an upfront investment: it slows down the first task to make the twentieth possible, and it has to be maintained like any other part of the system.

When we decide differently. For a throwaway exploration whose result nobody adopts, formal acceptance would be superfluous and the structure wasted.

4. Verifying: tests, CI and the gate before the effect

The same feedback that guides an agent cannot sign off on its work. A result from the environment tells it whether a step is good enough to take the next one — not whether the whole is something you can answer for. That is why the proof belongs outside: deterministic and outside the loop it verifies. It is the same construction rule with which we contain the probabilistic in production systems — applied to the way the software is made.

deterministic frame Plan Change Execution Feedback Verification Approval Task and test tests, types, CI human Effect

Diagram: the input is the task with its acceptance criteria. The loop may iterate — it only takes effect beyond verification and approval.

A gate is more than a test suite: tests verify behavior, a review verifies fit — whether a change belongs to this system and its intent. A change can pass every check and still not be accepted; this gap becomes visible as soon as agent-generated and human contributions run side by side. The objection that gates are bureaucracy misjudges them: whoever doesn't draw the line between a passed check and an accepted change hasn't abolished it but moved it behind the effect. Because a wrong inference in a multi-step loop carries through the following steps, the gate doesn't belong at the end but at every transition where an effect would occur.

Trade-off. Every gate buys reliability with part of the speed attributed to this way of working.

Cost. Gates are software themselves: they are designed, maintained and run on every change — and in our experience, a gate nobody trusts gets bypassed.

When we decide differently. Where a result only reaches the human as a suggestion anyway and doesn't bind them, they already are the gate; an additional one would be ceremony.

5. The verification burden and who carries it

Where more is generated, more has to be traced through. In a longitudinal survey of professional engineers — asked about assistive tools, not agentic work — the share of work shifts from producing to verifying and correcting, and a term of its own has been proposed there for this activity: supervisory engineering work. That it tends to increase rather than decrease with an agent's scope of action is our conclusion. It isn't distributed evenly in any case. An observational study of mature open-source projects — after the introduction of an assistive, not an agentic, tool — suggests that the additional verification and rework lands where judgment is scarcest: with the experienced people, who as a result review more and build less themselves. A tendency, not a law — for your own team, this can only be established, not assumed.

Which brings us to the most uncomfortable property of this way of working: perception is not a measuring instrument. In a controlled study, the participants' impression and the measurement diverged — it concerned assistive tools from an earlier generation of tooling, not agentic work, but this part of its finding doesn't depend on that. The participants were experienced and attentive; the misjudgment is therefore not carelessness but a property of the situation. Even where respondents consistently reported improvements, the same people simultaneously named deteriorations elsewhere. We are not saying anything about the direction of the effect, but this: your own impression cannot carry a decision about introduction, expansion or rollback. That requires measurement — whoever doesn't observe their system won't see this effect either. And the obvious conclusion that you need more experienced people is not the only one: in our assessment, a system with mechanically verifiable rules moves part of this judgment into the structure — and thereby lowers the demand this chapter describes.

Trade-off. Making verification capacity visible limits how much may run in parallel — it acts like a brake, and it is one.

Cost. Setting up measurement is work that delivers nothing in itself, and its results are often uncomfortable for those who drove the introduction.

When we decide differently. For a time-boxed trial with a clear stop condition, attentive observation is enough; formal measurement belongs to decisions that are meant to last.

6. Permissions, boundaries and human approval

Actions are not equivalent: reading, writing, executing, deploying and deleting hardly differ in effort, but differ considerably in the damage a single wrong inference can do. That is why trust is the wrong ordering criterion: the damage stays the same regardless of how reliably things went last time. Our criterion is reversibility; the following tiering is ours, not a standard:

ActionReversibilityApproval
Read, search, summarizecompletefree
Change in a dedicated working branchcompletefree
Execute in an isolated environmentcompletefree, limited to the environment
Write to a shared branchwith effortchecked automatically, then review
Deploy to productionlimitedhuman approval
Delete, migrate, change access or configurationnot at all or only at great costhuman approval, granted separately

As with probabilistic components in production, the same applies here: human control belongs neither everywhere nor nowhere — placed everywhere, it makes delegation pointless; placed nowhere, it moves the risk behind the effect. The objection that approvals make the approach uneconomical only hits the first variant — tiered by reversibility, most actions remain free. That agentic systems are treated in security work as a class of their own with their own standards, separate from the security of language-processing systems, is itself a hint: the difference lies in acting, not in answering.

Trade-off. Tiered approvals buy controllability with friction in everyday work — and in our experience, friction gets bypassed when it sits in the wrong place.

Cost. Every boundary is configuration that has to be maintained — an outdated approval list is more dangerous than none, because it feigns safety.

When we decide differently. Where an action can be undone without consequences, an approval would be pure ceremony — the yardstick is the possible damage, not the unease.

7. Why the same tools have opposite effects

The finding we rely on here is not about the tool but about the organization using it: AI tools amplify the strengths of well-set-up organizations and equally the dysfunctions of poorly set-up ones — surveyed for AI use in general, not for agentic work in particular. This becomes visible in a double movement: where more throughput is reported, more instability is reported at the same time. Whoever reads only one half has read only half the bill. The objection that in the end it is a tooling question fails precisely on this.

That tools amplify regardless of direction, we have described elsewhere; what counts here is what follows for this way of working. Where generating becomes cheaper, the direction of a piece of work is not set there but in defining — at exactly the station where the work is currently getting stuck anyway. Whoever is imprecise there multiplies the imprecision. On top of that comes a time profile that is reckoned with in practice and that can mislead an early assessment: a phase in which things initially get worse. Whoever assesses too early measures that phase and not what comes after — in whichever direction.

Trade-off. In our experience, investing in practice and structure has a stronger effect than the choice of tool, but it produces nothing to show off and takes longer.

Cost. Planning for the reported initial dip means having to defend it against expectations that promise immediate gains.

When we decide differently. Where practice, tests and delivery already hold up, the tooling question is the next open one — there, restraint is merely hesitation.

8. What changes in the role of experienced engineers

What recedes, according to what is reported for assistive tools, is writing itself; what increases is guiding, evaluating, correcting. We push back where this turns into a swan song: the developer doesn't disappear; their task moves up one level. What fills it — problem framing, architectural leadership, verification criteria, responsibility — is our description of that level, not an observed shift. And it makes judgment more expensive: whoever is responsible for a change they didn't write needs more contextual knowledge, not less, and has to be able to assess what they didn't think through themselves. Add to that the shape of the task itself: whoever verifies has to assess, in a short time, what came about over many steps.

This raises a question we cannot resolve. Until now, this judgment developed as a by-product of doing the work yourself: from mistakes, from detours, from the effort of having penetrated a system. Controlled learning studies — on learners, over hours, not on experienced professionals over years — report that assistance raises immediate performance while it can at the same time impair understanding, code reading and retention. One pointer in them is that forms of collaboration that require cognitive engagement preserved learning outcomes even with assistance; whether that holds equally over weeks in a real system, we don't know. Both are laboratory findings over short periods, and neither answers the question that has to be asked here: if experienced engineers implement less themselves — where will the practical experience come from in the future, from which judgment for architecture, review and responsibility grows?

Trade-off. Moving up a level means giving up the part of the work that used to fill most of the day — and whose side effects nobody has measured.

Cost. Being responsible for work you didn't write yourself is hard to make visible as an accomplishment — it leaves no artifact you can point to.

When we decide differently. Where signing off on a task takes longer than the task itself, delegation is the more expensive route.

9. Where agentic coding holds up — and where it doesn't

The first condition is the system. Where rules live in tests, types and contracts, the agent has feedback and the human has a gate; where they remain implicit, both are missing at once. The second condition concerns the expectation you start with. The same controlled study that already came up in the chapter on the verification burden did not find the expected speed-up; the authors themselves name familiarity with and maturity of the systems worked on as contributing factors, the sample was small, and they explicitly limit its generalizability. It concerned assistive tools from an earlier generation of tooling, not agentic work — it is no good as a verdict on agents. What it shows is something else and more timeless: a blanket speed-up narrative has failed to withstand measurement before. That is not an argument against the way of working, but one against assuming instead of measuring.

For us, what follows from this is not abstinence but restraint in making promises. In our experience, this way of working holds up for well-delimited tasks in systems with a safety net. It doesn't hold up for poorly specified problems, in poorly understood legacy systems without verifiability, and where the effect is not reversible. That is a judgment from experience, not an evidenced classification. The uncomfortable case remains the legacy system: there, the first sensible task may be to establish verifiability in the first place — under tighter approval, and as a project of its own, not on the side.

Trade-off. The verifiability condition hits first the systems where relief would be most urgent — there, the first task is not relief but verifiability.

Cost. Establishing verifiability after the fact is a project of its own with its own effort — it pays off, but not in the task that prompted it.

When we decide differently. Where a system is being replaced anyway, it doesn't pay to make it verifiable first just to let an agent work in it.

10. Common mistakes

The recurring patterns that make this way of working fail — almost all of them are variations on relieving one station and leaving the others unchanged:

  • Taking the perceived gain in generation for the whole calculation and overlooking the work it creates elsewhere.
  • Delegating a task for which nobody can say how “done” would be recognized.
  • Writing boundaries into the task description instead of into the structure of the system.
  • Placing the verification gate inside the same loop it is supposed to verify.
  • Mistaking a passed test for an accepted change.
  • Treating verification capacity as a by-product instead of a finite quantity — and then parallelizing.
  • Tiering approvals by trust in the agent instead of by the reversibility of the action.
  • Judging the effect by your own impression instead of measuring it.

11. Decision checklist

To clarify, in order, before and during the introduction of agentic work:

  • Task phrased verifiably? Is it established how “done” can be recognized — before anyone starts?
  • Boundaries in the system? Do module boundaries, types, contracts and tests carry the rules, or are they only in the description?
  • Gate outside the loop? Is the proof deterministic and independent of whatever produced it?
  • Test and acceptance separated? Is it clear that passed checks are not yet an accepted change?
  • Verification capacity planned? Is it planned as a finite quantity — and is it known on whom it lands?
  • Approvals by reversibility? Are reading, writing, executing, deploying and deleting treated differently?
  • Control where the effect occurs? Is there a human in the way where a wrong action triggers something — and only there?
  • Measured rather than felt? Is there a measurement of the effect, or does the participants' impression decide?
  • Condition before tool? Has it been checked whether practice, tests and delivery hold up — before the tool is supposed to be the answer?

Whoever can answer these questions has organized the work at the point where it gets stuck when generating becomes cheaper — and not where it used to get stuck.

FAQ

What distinguishes agentic coding from autocompletion and chat? In our reading, the scope of action, not the quality of the model. A completion suggests, a chat answers — in both cases a human executes and is thereby the reviewer of every step. An agent breaks down a goal, makes changes across several files, calls tools, reads their results and continues with them. What happens without a human intermediate step grows — and with it the amount of what has to be verified after the fact.

Does work disappear — or does it just shift? It shifts. This has been shown for assistive tools, not for agentic work: there, the share of producing falls and the share of verifying and reworking rises. That this shift tends to increase rather than decrease with an agent's scope of action is our conclusion, not a measurement. What used to happen on the side — deciding what should be built, and being certain that it is right — becomes the actual work.

What does a repository need for an agent to work in it reliably? We only work with mechanically verifiable rules: tests, types, contracts, clear module boundaries. At every step, they give the agent an answer as to whether it can stay on the path it has taken, and they give the human a point at which they can verify without having to reconstruct the entire sequence of changes in their head. Where these rules live only in a few people's heads, both are missing — the agent lacks support, the review lacks a yardstick.

How do you verify work you didn't write yourself? With a deterministic gate outside the loop it verifies: tests, type checking, CI, small reversible steps — and a review for fit, which tests don't capture. A passed test is not yet an accepted change. This includes measuring the overall effect instead of feeling it: where impression and measurement have diverged before, the impression cannot carry the decision about introduction or expansion.

Which actions should an agent never perform without human approval? Our rule: those that are only partially reversible or not reversible at all. Deploying to production; deleting, migrating, changing access or configuration. Reading, changing things in a dedicated working branch and executing in an isolated environment remain free. The yardstick is the reversibility of the action, not trust in the agent — because the damage stays the same regardless of how reliably things went last time. In our experience, approvals placed everywhere get bypassed in everyday work.

Where is agentic coding the wrong choice? In our experience, where verifiability is missing: in poorly understood legacy systems without tests and contracts, with poorly specified problems whose acceptance nobody can formulate, and wherever an action is not reversible. That is a judgment from experience, not an evidenced classification. In a legacy system, the first sensible task may be to establish verifiability in the first place — as a project of its own and under tighter approval, not on the side.

Further reading

The foundation is the Batunet Engineering Method: frame the problem before building; prove instead of assuming; keep judgment with humans.

Closing engineering principle

What makes a system hold up was never the typing. It was clarity about what should be built, proof that it is right, and the willingness to stand behind it. As long as generating was expensive, these three things could be taken care of on the side — they had no place of their own in the working day, because generating filled it. Where that no longer holds, they need one. That is the change, and it is more uncomfortable than the promise it comes with: the work doesn't get smaller — the point where it gets stuck changes, and the new one is harder to delegate than the old. Whoever wants to hand off a task has to know it well enough to say how its success can be recognized. That has never been busywork.


An agent can carry out a task. Only someone who could say what needed to be done — and can judge whether it was done — can take responsibility for it.

Referenced entities

Knowledge graph

Continue your engineering journey.

Related concepts, decisions, playbooks and perspectives — as one connected path, not a list of links.

A concrete project in this field?

Reference Guides show how we think. For your system, talk to our management — technical, no sales pitch.