Data Modeling That Lasts
Of everything that makes up a system, the data model lives the longest. Frameworks come and go, interfaces get replaced, even the database can be swapped out — the shape of the data remains and carries everything above it. How to design it so that it lasts a decade. A decision document for CTOs, architects and senior engineers.
What is this? · Reference Guide
A solid guide to an engineering question — with trade-offs, costs and the case in which we decide differently. Not an opinion piece, but a reference text. Go to overview
- Author
- Batunet Engineering
- Reading time
- 18 min
- Level
- In depth
- Status
- Approved
- Last reviewed
- 21 July 2026
- Updated
- 21 July 2026
On this page
On this page
- 1. Why the data model outlives everything
- 2. Model the domain, not the interface
- 3. Modeling for change
- 4. Identity and relationships
- 5. Normalization and its pragmatic limits
- 6. The cost of a wrong model
- 7. Evolving the model
- 8. Where model and code drift apart
- 9. Common mistakes
- 10. Decision checklist
- FAQ
- Further reading
- Closing engineering principle
Look closely at a long-lived system and you notice an order of permanence. The outermost layer — the user interface — changes most often. Beneath it sits the framework, which gets upgraded or replaced over the years. Deeper still, the database, which can be replaced with effort. And at the very bottom, the most permanent of all, lies the shape of the data itself: which things exist, how they relate, what distinguishes them. That shape outlives everything above it — and that is why it determines longevity more than any other choice.
This document treats the data model as what it is: the foundation a system stands on. It is deliberately database- and framework-neutral, because it is about the structure of the data, not the technology that stores it. Other texts in this collection keep pointing to the data model as the layer that really matters; this is the text that spells it out. No figures are given.
1. Why the data model outlives everything
A data model outlives the technology around it because it does not represent the technology — it represents the reality the system manages. A customer, an order, a contract, a case — these things and their relationships change far more slowly than the frameworks used to process them. You can rebuild the interface, switch frameworks, even migrate the database, without changing what a customer is or how it relates to an order. That is precisely where the model's permanence comes from: it is indifferent to the technology.
This dictates where care belongs. Because the data model lives the longest and carries everything above it, it deserves the most attention during design — more than the choice of framework, more than the choice of database, more than the design of the interface. A mistake in the interface costs a rebuild; a mistake in the data model costs a rebuild of everything that stands on it. Teams that reverse this order — arguing at length about the framework while the model takes shape on the side — have put their effort in the wrong place.
Diagram: The deeper the layer, the longer it lives. The data model at the bottom outlasts interface, framework and database — and carries them all.
Trade-off. Giving the data model the most care means investing more time in something that stays invisible at first glance — people see the interface, not the model underneath.
Cost. A well-thought-out model takes longer to create than a hastily sketched table structure, and it requires experience you cannot skip.
When we decide otherwise. For a throwaway prototype whose data will never be carried over into a long-lived system, full rigor is overkill — speed matters there, and the model may stay rough.
2. Model the domain, not the interface
The most common and most consequential mistake is shaping the data around whatever the screen or the API needs right now. A form has certain fields, so you create a table with exactly those fields; a view shows a certain combination, so you store the data in that combination. It is tempting because it is the least work in the short term — and expensive in the long term, because interfaces change and then the model no longer fits.
The data model should represent the domain, not the interface: things as they actually are and relate in reality, regardless of how they are displayed or queried today. A model that captures the domain survives every redesign of the interface, because the domain is stable while the view changes. A model that mirrors the interface has to move along with every change to it — and becomes an obstacle at the first major redesign. The question is not "which fields does the screen show?" but "what is this thing really, and what does it relate to?".
Trade-off. Modeling the domain instead of the interface takes more thought up front — you have to understand what a thing really is instead of just copying what the form shows.
Cost. A domain-oriented model sometimes requires a translation between the shape of the data and the shape of the display that you would otherwise have avoided.
When we decide otherwise. Where a display and the domain truly coincide and will permanently stay that way, the separation is overengineered; you simply model what is both at once anyway.
3. Modeling for change
A data model meant to carry a system for a decade has to be changeable, because the domain grows and shifts over time. As with any long-lived structure, the art is not predicting the future but laying out the model so that change stays cheap. Above all, that means being able to grow additively. A new feature arrives as a new field or a new, connected thing, without shifting the meaning of what already exists. A model that can only grow by restructuring what is there becomes riskier with every change.
| Technique | What it achieves | Its price |
|---|---|---|
| Extend additively | New things are added without shifting existing ones | Model grows and needs to stay orderly |
| Restraint | only commit to what the domain requires | initially less tightly tailored to today |
| Old and new in parallel | Change becomes reversible, without downtime | two shapes to maintain for a while |
| Room for the unclear | the one unforeseen requirement doesn't break it | less elegance in the moment |
Part of this is not committing to more than you know. A model that crams every conceivable property into a rigid structure early on is brittle against the one property nobody thought of. Restraint in the model — committing only to what the domain really demands and leaving room for what is not yet clear — is the same stance as reversibility in architecture: keep the reversible cheap, commit to the irreversible late and deliberately. Because a change to the data model of a running system is among the most irreversible interventions there are — it touches every existing record.
Trade-off. Modeling for change means committing to less and staying more general — the model is initially less tightly tailored to today's case.
Cost. A changeable model demands discipline with every extension so that additive growth does not tip into sprawl, plus the willingness to fix structure only once it is clear.
When we decide otherwise. Where a structure is fixed by law or by an immutable external contract, you define it — firmly and in detail; openness toward something unchangeable would be wasted.
4. Identity and relationships
Two decisions in the data model are more permanent than all others: what makes a thing unique, and how things relate to each other. Identity — what makes a record exactly this one and distinguishes it from all others — is the foundation on which every relationship, every reference, every merge rests. A poorly chosen identity — one that can change, that is not truly unique, that mixes something domain-specific with something technical — poisons the whole model, because everything builds on it. Identity is the decision that is hardest to take back.
Diagram: Identity makes each thing distinguishable; the relationships say how the domain works. Both last the longest and are the most expensive to change.
Relationships are the second permanent layer: which things belong to which, whether one relates to many or many to many, whether a connection is mandatory or optional. This structure forms the model's actual statement about how the domain works — and changing it almost always means restructuring existing data, not just writing new code. That is why identity and relationships deserve the greatest care in the entire design: they are what lasts longest and is most expensive to correct. Get them right and the model is right at its core; get them wrong and you carry the mistake through its entire lifetime.
Trade-off. Choosing identity and relationships carefully takes the most thinking time in the design — in exchange for the most permanent part of the model holding up from day one.
Cost. This care requires genuinely understanding the domain before committing — work you cannot shortcut without paying dearly to make up for it later.
When we decide otherwise. Where a thing is obviously and stably identified and its relationships are unambiguous, you decide quickly; full rigor is reserved for cases where identity or relationship is not clear at first glance.
5. Normalization and its pragmatic limits
The classic discipline of data modeling is normalization: store every fact exactly once, in one authoritative place, so that contradictory copies cannot exist. It is the same idea as "one piece of knowledge, one representation" — applied to data. A normalized model is truthful: it cannot contradict itself, because every fact has only one place. For correctness over the years, that is worth a great deal, since contradictory data is one of the most stubborn sources of errors there is.
Like every principle, this one has its limits. Strict normalization can make queries cumbersome and expensive, and there are cases where deliberate, controlled redundancy is the better choice — as long as you know you are taking it on and make sure the copies stay consistent. The art is neither to normalize dogmatically nor to simplify dogmatically, but to ask: is this fact authoritative here, or just a copy? And if it is a copy: who keeps it consistent? A model that answers these questions deliberately is durable; one that accumulates redundancy unnoticed drifts into contradictions over time.
| Approach | Strength | Price |
|---|---|---|
| Strictly normalized | no contradictions, every fact stored once | queries more cumbersome, more joining |
| Deliberately redundant | queries simpler, faster | copies must be kept consistent |
| Unwittingly redundant | — | drifts into contradictions, the most expensive case |
Trade-off. Strict normalization buys freedom from contradictions at the cost of more cumbersome queries; deliberate redundancy buys simplicity at the cost of the obligation to keep copies consistent.
Cost. Both deliberate paths require making the decision and documenting it; the expensive case is the third, unwitting variant, in which redundancy builds up unnoticed.
When we decide otherwise. Where correctness trumps everything — money, legal matters, inventory — you normalize strictly; where a measured read load makes joining too expensive, you accept controlled redundancy and name who keeps it consistent.
6. The cost of a wrong model
A wrong data model is the most expensive mistake a system can carry, because everything builds on it. A mistake in the interface is local; a mistake in the business logic is contained; a mistake in the data model is everywhere, because every line of code that works with the data has assumed the wrong structure. Worse still: the wrong model fills up with data. Every day the system runs, the body of data in the wrong shape grows — data you cannot throw away when you correct the model, but have to migrate.
This explains why a data model deserves so much care up front: correcting it is not just a code change but a migration of every existing record, often while the system is running — one of the most demanding and risky operations there is. A model gotten right early spares you that operation; one gotten wrong early forces it eventually, at a point when the data volume is large and the system is critical. The care you invest at the start is insurance against the most expensive repair a system knows.
Trade-off. Taking the high cost of a wrong model seriously means being slower at the start — in exchange for the certainty of avoiding the far more expensive correction later.
Cost. The initial care competes with the pressure to deliver something visible quickly, and it only pays off over the years.
When we decide otherwise. For a system that is foreseeably short-lived and whose data will never need to be migrated, full caution is overkill; its value rises with the expected lifetime and the criticality of the data.
7. Evolving the model
However permanent a data model is, it is not immutable — and the ability to evolve it safely is part of its longevity. The discipline for this is the same one that makes even the riskiest change manageable: small, reversible steps, additive rather than restructuring existing data, old and new side by side for a while. You add the new shape, migrate the existing data with verification, switch the code over and remove the old shape only once nothing depends on it anymore. That way a model change remains an orderly process rather than a leap into the unknown — carried out while the system is running, without downtime.
This evolution matters so much that it should shape the design: a model that is easy to extend additively is superior to one that turns every change into a restructuring — even if both do the same job today. With the data model, too, changeability is the real yardstick, not elegance in the moment. The concrete mechanics of this verified, reversible migration are covered in detail in a separate text; for the design, what counts is laying out the model from the start so that those mechanics can apply at all.
Trade-off. Designing for evolution means favoring additively extensible structures, which are not always the most compact solution for today.
Cost. Even with the best preparation, a model change remains work — migration, verification, parallel operation — that you plan for instead of underestimating.
When we decide otherwise. Where a model will almost certainly never be extended, it may be more compact and narrower; provisioning for change pays off where the domain is foreseeably going to grow.
8. Where model and code drift apart
A data model lives in two places at once: as a structure in the database and as a representation in the code that works with it. Both must say the same thing about the same matter — and because they can be changed independently, their drifting apart is a recurring, insidious source of bugs. If one side changes without the other following, errors arise whose effects show up far from their cause and are therefore hard to find. A durable data model is of little use if its representation in code silently deviates from it.
The consequence for the design is to keep this coupling visible and verifiable instead of relying on the attentiveness of individuals: a mismatch between the structure and its representation should surface early and loudly, not late as a puzzling symptom. The data model is therefore not just a question of design but also of operational discipline — keeping the two representations of the same model in sync. Taking this silent coupling seriously spares you an entire class of stubborn bugs.
Trade-off. Making the coupling visible costs an additional safeguard that needs maintaining — in exchange for a drift surfacing early and close to its cause.
Cost. Every check is effort, and it makes the moment of changing the structure or its representation somewhat more cumbersome, because now both have to match.
When we decide otherwise. In a very small system whose model one person can easily keep in view, attentiveness is enough; the safeguard pays off as soon as many hands work on the same model.
9. Common mistakes
The recurring patterns that make data models fail — almost all of them are variations of treating the permanent as incidental:
- Shaping the model around the interface instead of the domain — and discovering at the first redesign that it no longer fits.
- Giving the framework and the database more care than the model that will outlive them both.
- Choosing identity poorly — mutable, not truly unique, mixing domain concerns with technical ones — and carrying the mistake through the entire model.
- Defining relationships imprecisely and paying for their later correction with an expensive restructuring of existing data.
- Accumulating redundancy unnoticed until the model drifts into contradictions.
- Committing to too much too early and being brittle against the one requirement nobody thought of.
- Laying out the model so that every change is a restructuring, instead of being able to extend it additively.
- Letting the structure and its representation in code drift apart because nobody made the silent coupling visible.
- Underestimating the cost of a wrong model and skimping on care at the start, where it would be most valuable.
10. Decision checklist
To clarify in order, before and during the design of a data model:
- Domain or interface? Does the model represent what things really are — or just what a screen happens to show?
- Care distributed correctly? Does the model get more attention than the framework and database it will outlive?
- Identity sound? Is every thing identified by something stable and truly unique that does not mix domain concerns with technical ones?
- Relationships clear? Is it defined which things relate to each other and how — and is that faithful to the domain, not to the current view?
- Additively extensible? Can the model grow without shifting the meaning of what already exists?
- Redundancy deliberate? Is every copy of a fact intentional, with a named owner who keeps it consistent?
- Evolution considered? Can the model be changed in small, reversible steps while the system is running?
- Coupling visible? Does a drift between the structure and its representation in code surface early and close to its cause?
If you can answer these questions, you have designed a model that holds up — not one that gives way at the first change.
FAQ
Why is the data model more important than the choice of database or framework? Because it outlives both. You can upgrade the framework, rebuild the interface, even migrate the database, without changing what a customer is or how it relates to an order. The shape of the data remains and carries everything above it — which is why it determines longevity more than any technology choice.
What does "model the domain instead of the interface" mean? Shaping the data around what things really are and how they relate, not around what a form or a view happens to show. A domain-oriented model survives every redesign of the interface because the domain is stable; an interface-oriented model has to move along with every display change and soon becomes an obstacle.
Why are identity and relationships so critical? Because everything builds on them and they are the most expensive to correct. A poorly chosen identity — mutable or not truly unique — poisons every reference that rests on it. Changing a wrong relationship almost always means restructuring existing data. These decisions last the longest and are the least forgiving.
Should you always normalize strictly? As a default stance, yes, because a normalized model cannot contradict itself. But it has its limits: where a measured read load makes joining too expensive, deliberate, controlled redundancy is justifiable — as long as you know you are taking it on and name who keeps the copies consistent. The expensive case is unwitting redundancy that drifts into contradictions.
What makes a wrong data model so expensive? That everything builds on it and that it fills up with data. A modeling mistake affects every line of code that works with the data, and every correction is not just a code change but the migration of every existing record — often while the system is running. It is the most expensive repair a system knows, and care at the start is the insurance against it.
How do you change a data model safely? In small, reversible steps: add the new shape additively, migrate the existing data with verification, switch the code over, and remove the old shape only once nothing depends on it anymore — while the system is running, without downtime. The same discipline that makes database migrations manageable applies to every change to the model.
Further reading
- Software that still runs in ten years — why longevity is the fundamental question, whose deepest layer is the data model.
- PostgreSQL or MySQL for long-lived systems — why the model matters more than the choice of database.
- Zero-downtime database migrations — the mechanics of evolving a model reversibly while the system is running.
- When schema and model drifted apart — what happens when the structure and its representation silently diverge.
It is grounded in the Batunet Engineering Method: model the domain, build for change, give the most care to what lasts longest.
Closing engineering principle
Whoever builds a system for a decade builds its data model first — because it is the only thing that will almost certainly still be there in a decade. Everything above it will be replaced: the interface several times, the framework perhaps, the database possibly. The shape of the data survives them all, because it does not represent the technology but the reality the system manages. That is why data modeling is not one step among many, but the decision that carries the deepest and lasts the longest. You often recognize good engineers not by the interface they build, but by the model underneath it that still holds when the interface has long since been replaced three times.
The interface is what you see; the data model is what remains. You build a long-lived system from the bottom up — and at the very bottom lies the shape of the data.
Referenced entities
Continue your engineering journey.
Related concepts, decisions, playbooks and perspectives — as one connected path, not a list of links.
Services
Concepts
A concrete project in this field?
Reference Guides show how we think. For your system, talk to our management — technical, no sales pitch.
