Reference Guide · Architecture

Legacy Modernization Without a Big Bang

How to replace a legacy system in production without stopping it — incrementally, reversibly, under observation. For CTOs, technical directors and architects.

What is this? · Reference Guide

A solid guide to an engineering question — with trade-offs, costs and the case in which we decide differently. Not an opinion piece, but a reference text. Go to overview

Author
Batunet Engineering
Reading time
15 min
Level
In depth
Status
Approved
Last reviewed
21 July 2026
Updated
21 July 2026
On this page

Almost every company that has been running software long enough reaches the same point: a system carries the business, but nobody likes changing it anymore. It has become opaque, the people who built it are gone, and every adjustment feels like open-heart surgery. The obvious wish is to rebuild it properly, once and for all.

That wish is the most expensive mistake in software engineering. A legacy system is not a design you cleanly replace — it is the frozen state of years of decisions, special cases and silent requirements, most of which are written down nowhere. This text is about how to replace such a system without the big bang: incrementally, in small reversible steps, with the old system as a safety net until the new one can carry the load.


Why the big rewrite fails

The big-bang rewrite promises to get everything right at once: build the new system, migrate everything, switch over on a cutover date. It fails for three structural reasons, not for lack of skill.

First: the old system contains knowledge nobody has anymore. The thousand small special cases that have accumulated over the years are requirements — just invisible ones. A greenfield rebuild does not reproduce them; it discovers them in production, one outage at a time.

Second: while you rebuild, the old system does not stand still. The business still needs changes. Either you freeze development on the legacy system — and the company loses its ability to move for the duration of the rebuild — or you maintain both systems in parallel and chase a moving target.

Third: the cutover date concentrates the entire risk into a single moment. Everything has to work at the same time — code, data, integrations, operations. If something fails, there is no partial way back, only a return to square one. That is the opposite of how we build: in small, reversible steps.

The recommendation follows from this: break the risk down into small pieces that can each be handled on their own, instead of concentrating it on one date. The price is that the replacement takes longer and two systems have to coexist the whole time — more operational effort, a temporary seam you maintain yourself. We decide otherwise only when the system is small enough to be rebuilt completely within a few weeks with a real test safety net — then a clean rebuild can be cheaper than the seam.

What "legacy" really means

Legacy does not mean "old". A twenty-year-old system that is understood, tested and maintained is not a legacy problem — it is mature technology. Legacy is a system that carries the business, runs in production and is poorly understood. The value lies in "carries the business"; the risk lies in "poorly understood".

This distinction determines the entire strategy. The goal of modernization is not to replace old code with new code. The goal is to regain understanding and changeability without destroying the value contained in what exists. Forget that, and you optimize for new technology and risk exactly the one thing that makes the system valuable: that it works.

The principle: understand, contain, replace

Incremental replacement always follows the same order. First understand what a part does — observably, not by assumption. Then contain it: lay a seam between old and new behind which you can work without touching the rest. Then replace, piece by piece, putting each replaced piece into production immediately.

The core is reversibility. Every step is cut so that it can be shipped individually and rolled back individually. You never bet the entire system on one change; you move a thin slice, check whether it holds and move on to the next. If a slice goes wrong, the damage is limited to that slice — not to the cutover date.

MoveWhat it doesResult
Understandmake the actual behavior observableevidence-based knowledge instead of assumption
Containlay a seam between old and newworking space without touching the rest
Replacemove slice by slice, into production immediatelythe old shrinks, the new grows

The strangler fig approach

The load-bearing pattern is called strangler fig, after the fig that grows around a tree and eventually replaces it while the tree stays standing the whole time. You put a facade in front of the legacy system — the point where all requests arrive. Initially, the facade passes everything through to the old system unchanged. Then you move one business slice after another into the new system and have the facade reroute exactly those slices. The old shrinks, the new grows, and at no point does the system stand still.

StateFacadeLegacy systemNew
Startpasses everything throughcarries everythingempty
Transitionreroutes individual slicesshrinksgrows
Endpoints only to the newswitched offcarries everything
Client Facade (the seam) Legacy system being retired New slice by slice

Diagram: The facade is the only fixed point; behind it, the system migrates.

The recommendation: introduce the facade first and ship it in pass-through mode before you replace anything. That way you separate the risk of the seam from the risk of the migration. The price is an additional layer every request passes through — some latency, a piece of infrastructure you have to operate and secure, and the discipline that no new feature gets connected directly to the legacy system, bypassing the facade. We decide otherwise when the system has no clear entry point where a facade can be placed — a thick desktop application without a service boundary, for instance; there the seam first has to be created through an interface before the replacement can begin.

The seam: facade and anti-corruption layer

The seam has two jobs. As a facade, it decides which request goes to the old system and which to the new. As an anti-corruption layer, it translates between the old and the new world so that the concepts and errors of the legacy system do not seep into the new domain. Without this translation, the new system inherits the old one's modeling mistakes — and you have invested a lot of work just to repackage the same problems.

The recommendation: keep the new domain independent of the old model; everything coming from the legacy system is translated at the seam into the language of the new domain. The price is translation code that looks duplicated and has to be maintained as long as both worlds coexist. We decide otherwise when the old model is demonstrably clean and is meant to live on — then the translation is unnecessary friction, and you adopt the model deliberately instead of encapsulating it.

Data is the real problem

Code can be moved slice by slice. Data, not so easily: it is shared, it is large, and the database outlives every application layer above it. The hardest part of any modernization is almost never the code — it is moving the data without losing it and without stopping the system.

The tool for this is the step-by-step, reversible schema change — expand, migrate, contract. First expand: add the new schema additively alongside the old one, without removing anything. Then migrate: write to new and old in parallel for a while (dual write), backfill existing data in the background, gradually switch reads over to the new. Only when nothing reads the old anymore do you contract: remove the old field. Every single change is backward-compatible on its own and therefore reversible.

Expand new alongside old Migrate write in parallel, backfill Contract remove old

Diagram: No field is changed; one is added, filled, then the old one is removed.

The recommendation: migrate data early and in parallel, not as the last step before the cutover date. The price is a dual-write phase in which two sources have to be kept consistent, plus the effort of backfilling existing data and verifying that both sides match. We decide otherwise for small, non-critical data volumes with an acceptable maintenance window — there a one-time migration within the window is simpler than the ongoing dual-write machinery.

Branch by abstraction and parallel operation

Not every part sits comfortably behind the facade. A component wired deep into the system — a calculation, access to an external service — can be replaced with branch by abstraction: you place an abstraction over the old implementation, write the new one behind the same abstraction, switch over via a toggle and finally remove the old one. The rebuild happens in the live code, without a long-lived side branch.

Before switching over comes parallel operation: run the new and old implementations together for a while, compare both results, but keep treating only the old one as authoritative. Only when the new one delivers the same results across real traffic do you switch over. That way correctness is proven in production before it takes on responsibility — not afterwards.

The recommendation: prove the new implementation in parallel operation against real traffic before it becomes authoritative. The price is double execution for the duration of the comparison — more load, more instrumentation, and dealing with discrepancies, which often uncover forgotten special cases. We decide otherwise where double execution would have side effects that cannot safely be duplicated — payments or shipping, for instance; there you verify against recorded traffic in a shadow environment instead of live.

Order: riskiest first, in thin slices

Where to start? Not with the easiest part, to have something to show quickly, and not with the biggest, to get it over with. The riskiest, least understood path comes first — but as the thinnest possible slice that runs end to end through the seam. That way you clear the biggest unknown at the start, while change is still cheap, instead of discovering it at the end.

A good first slice is self-contained in business terms, small and genuinely in production. It proves the whole chain — facade, new domain, data, operations — on a real but limited case. What it teaches you about the seam and the data paths carries every subsequent slice.

The recommendation: cut the work along business capabilities, not technical layers, and begin with the riskiest thin slice. The price is that the first slice seems disproportionately expensive, because it builds the entire seam along with it before visible progress appears. We decide otherwise when a part is under acute pressure — a security risk, a data protection problem, a component that keeps failing; then that part moves to the front regardless of its risk profile.

Characterization tests: the safety net before the rebuild

You cannot safely change what you cannot observe. Before a slice moves, you capture the current behavior of the legacy system in characterization tests — tests that describe not what the system should do, but what it does, including its quirks. They are the safety net: if the new behavior deviates, they fire before production does.

The decisive point is order and separation. First the behavior is put under the safety net, then it is moved, and during the move the behavior is not improved at the same time. A modernization that builds in new features simultaneously can no longer interpret discrepancies — every difference could be intentional or a bug.

The recommendation: write characterization tests before the rebuild and keep migration and feature changes strictly separate. The price is that you write tests for behavior you intend to replace anyway, and that desired improvements have to wait until the move is done. We decide otherwise where the old behavior is demonstrably wrong and nobody relies on it — then you don't freeze the bug in the test but correct it deliberately and document the deviation.

Observability and a ticket back

Incremental replacement depends on two things in operation: you have to see what is happening, and you have to be able to go back. Every slice goes live behind a toggle, with observability from the first minute — old and new paths deliver comparable signals, so that a discrepancy is visible immediately. The old path stays operational until the new slice has built up trust across enough real traffic.

The recommendation: put every slice behind a toggle that lets you switch back to the old path within seconds, and keep the old path until trust is established. The price is that both paths have to remain runnable for a while — double code, double operations, and the obligation to actually remove the old path at the end instead of letting it creep into a permanent state. We decide otherwise for a slice with no significant risk whose way back costs more than its potential damage — there you switch over directly and save yourself the double path.

When a rewrite is the right call after all

The honest answer is part of giving advice, even when it argues against your own method. There are cases where the incremental path is the wrong one. When the platform itself is dead — a runtime without security updates, for which there are no people and no future anymore — no rebuild within the existing system will help. When the business problem has changed so fundamentally that the old model no longer answers the right question, you don't migrate a wrong model; you build the right one. And when the system is small enough to be rebuilt completely in a short time with a real test safety net, the rebuild can be cheaper than the seam the replacement would need.

The common denominator: a rewrite is right when the risk is small and manageable — not when frustration with the old system is high. The question is never "do we want it new?", but "can we take the same responsibility for the rebuild as for the existing system?". Where the answer is no, the incremental path applies.

Common mistakes

The same patterns keep making modernizations expensive or dangerous:

  • The big-bang cutover date that concentrates the entire risk into one moment with no partial way back.
  • The greenfield rewrite that rediscovers the old system's invisible requirements in production.
  • The feature freeze for the entire rebuild period, which robs the business of its agility and builds up pressure that eats away at diligence.
  • Migration and new features at the same time, so that discrepancies can no longer be interpreted.
  • Data last, instead of early and in parallel — the part most likely to tip over.
  • No seam: new features are attached directly to the legacy system and extend exactly what you wanted to replace.
  • The old path that stays "temporarily" after the switchover and becomes a permanent second system.

Checklist

Questions a CTO can ask of a modernization project. They are diagnostic questions, not verdicts.

  • Is there a seam — a place where we can separate old and new? Without a seam, there is no incremental replacement.
  • Is the current behavior under a test safety net before we touch it? What you cannot observe, you cannot safely change.
  • Can every slice be shipped individually and rolled back individually? Otherwise it is a big bang in installments.
  • Are we migrating data early and backward-compatibly, or saving it for the end? Data tips over last and most expensively.
  • Are we cleanly separating migration and feature changes? Doing both at once makes discrepancies unreadable.
  • Does the old path stay operational until trust is established — and is there a plan to remove it afterwards? A way back that is never closed becomes a second system.
  • Are we starting with the riskiest part in a thin slice, instead of the most visible one? Risk belongs at the beginning.
  • Could we take the same responsibility for a rebuild as for the existing system? If not, the incremental path is the right one.

FAQ

Doesn't the incremental path take longer than a rewrite? Up to the first visible slice, often yes, because the seam is built along with it. Over the whole project, usually not — and above all, the business keeps running the entire time. The rewrite only seems faster because its risk stays invisible until the cutover date cashes it in.

Isn't an additional facade itself new complexity? Yes, and it is temporary. The facade is the price of reversibility: it lets you move slice by slice without touching the rest. It is removed once the old system is retired. If it remains permanently without a purpose, it was in the wrong place.

What if nobody knows anymore what the legacy system actually does? Then understanding is the first step, not building. Characterization tests and observability make the actual behavior visible before it moves. You uncover the knowledge instead of guessing it — exactly the step the big-bang rewrite skips.

Can we ship new features during the modernization? Yes — that is one of the main reasons for the incremental path. But not in the same slice as a move. New features are built in the new system behind the seam; moving and changing functionality stay separate, so that discrepancies remain interpretable.

How do we know we are done? When nothing goes through the facade to the legacy system anymore, no source reads the old data schema anymore and the old path has been removed. Done is not the last piece of new code, but the moment the old system can be switched off without anyone noticing.

Where do you start — which slice first? With the riskiest one, in the thinnest possible form. Moving the risky part first uncovers the hard problems while the project is still young and you can turn back — not at the end, when everything depends on it. The slice stays thin so that a failure is cheap. Moving the risky part late and thick is the most common way to squander the incremental advantage.

Further reading

It is grounded in the Batunet Engineering Method: understand, decide, prove, build, harden, operate — in small, reversible steps.


A successful modernization is invisible. There is no cutover date, no anxious waiting, no night when everything is at stake. At some point the new system is running, the old one is off, and nobody can name the moment it happened.

Referenced entities

A concrete project in this field?

Reference Guides show how we think. For your system, talk to our management — technical, no sales pitch.