Caching Is Not a Performance Feature
Caching is seen as the fastest remedy for a slow system. That is exactly where the misunderstanding lies: a cache is not an optimization you bolt on afterwards, but a second source of truth — with everything that entails. For CTOs, architects, tech leads and senior engineers.
What is this? · Reference Guide
A solid guide to an engineering question — with trade-offs, costs and the case in which we decide differently. Not an opinion piece, but a reference text. Go to overview
- Author
- Batunet Engineering
- Reading time
- 19 min
- Level
- In depth
- Status
- Approved
- Last reviewed
- 21 July 2026
- Updated
- 21 July 2026
On this page
When a system gets slow, the phrase comes up early: "Let's just cache it." It sounds like a small, local measure — a few lines, a key, an expiry. In reality, at that moment you're making one of the most consequential architectural decisions there is: you're introducing a second copy of the truth and committing to keeping it in sync with the first. Everything that gets hard afterwards — wrong data, mysterious bugs, outages under load — originates here.
This text treats caching as what it is: a decision about consistency, not about speed. It is deliberately framework-, database- and cloud-neutral. The principles apply to a browser cache just as much as to a CDN, an in-process in-memory cache or a distributed cache across many nodes.
1. Why caching is misunderstood
The misunderstanding starts with the category. Caching gets filed under "performance," alongside indexes, compression and faster servers. But those measures only change how quickly the same truth is delivered. A cache changes something else: it delivers a possibly outdated truth — quickly. The gain in latency is paid for with a loss of freshness. That isn't an optimization in the strict sense; it's a trade.
The reason caching looks so tempting is real: most systems read far more often than they write, and the same data is requested again and again. Computing a value once and serving it many times saves real work. The problem isn't that caching doesn't work — it works almost too well. It hides slowness so reliably that it also masks structural problems you actually ought to fix.
This leads to a stance that underpins this entire text: a cache is the answer to a measured problem, not a reflex to a suspected one. Before you cache, you have to know what is slow and why. Often the cause is a missing database index, an N+1 query or a call that happens synchronously but belongs in an asynchronous path. In those cases, the cache treats the symptom and preserves the disease.
Trade-off. A cache trades consistency and simplicity for latency and relief of the source. Both are real; neither is free.
Cost. Every cached value doubles the number of places where a truth exists. Someone has to maintain that duplication — in the code, in their head and in operations. The real cost center isn't the memory, but the ongoing obligation to invalidate.
When we decide differently. If a problem can be solved cleanly at the source — a missing index, a decoupled computation, a better query — we solve it there and don't cache. We only reach for a cache once the source is as fast as it can be and the remaining load comes from sheer repetition.
Diagram: the cache sits between request and origin. From the moment it holds a copy, the two can drift apart — that is the core of every caching question.
2. What should never be cached
Before the question of how to cache comes the question of what you may cache at all. Some data tolerates aging, other data doesn't. The line runs where an outdated answer isn't just unattractive, but wrong or dangerous.
Authorization and permission decisions don't belong in a cache. Whether someone is allowed to do something has to hold at the moment of the action, not at the moment of an earlier call. A cache that says "allowed" after the right has been revoked is a security hole, not a performance gain. Nor do personal or tenant-bound data belong in a shared cache without strict separation by identity — the most common reason one user sees another user's data is a cache key that doesn't include the identity.
Values that must be correct at the moment they are read don't belong in a cache either: the current account balance before a debit, the available stock in the last step of an order, a price that is legally binding once quoted. For such values, "almost current" is the same as "wrong."
Trade-off. The temptation is to cache these values too, because they are often hot. But the trade would be latency for correctness — and correctness is non-negotiable where money, law or access are involved.
Cost. Keeping this data fast at the source takes real work: indexes, lean queries, possibly a dedicated read model. We bear these costs deliberately rather than shifting them into a cache that jeopardizes correctness.
When we decide differently. You can cache the derivation of such values without caching the value itself — for example, an expensive aggregation whose result is still checked against the live state at the end. Authority then stays with the source, and the cache only carries the groundwork.
3. Cache invalidation
The well-known saying that there are only two hard problems in computer science — cache invalidation and naming things — is half serious. Invalidation is hard because it requires a system to know when a copy has become invalid even though the change happened somewhere else entirely. At its core, there are three ways to organize that knowledge.
The first is expiry: you trust the copy for a fixed duration and discard it afterwards. The second is active invalidation: whoever changes the truth deliberately deletes or overwrites the affected copy. The third is event-driven invalidation: changes produce events, and the cache reacts to them. The three aren't mutually exclusive; robust systems combine them.
| Strategy | Consistency | Effort | Main risk | Fits when |
|---|---|---|---|---|
| Expiry (TTL) | weak, delayed | low | staleness within the window | aging is tolerable |
| Active invalidation | strong on every write path | high | forgotten write path | write paths are few and known |
| Event-driven | strong, decoupled | high, infrastructure | lost/delayed events | many writers, clear events |
The hardest part of active invalidation isn't the deletion itself but completeness: every path through which the truth can change must invalidate the copy. If a write path is forgotten — an import, an admin tool, a background job — a wrong copy remains, and the bug only shows up sporadically. That is exactly why expiry is so widespread despite its weakness: it is the only strategy that still works when you've overlooked a write path.
Trade-off. Active and event-driven invalidation buy freshness with coupling and an obligation to completeness. Expiry buys simplicity with guaranteed but bounded staleness.
Cost. The most expensive variant is event-driven: it requires reliable event distribution and with it some of the discipline that distributed message processing also needs — including the question of what happens when an event is lost.
When we decide differently. As a baseline we choose expiry, because it is safe even with incomplete knowledge. We add active invalidation where write paths are manageable and staleness within the expiry window is a noticeable problem. We only go event-driven once there are many independent writers and an event infrastructure already exists anyway.
4. Cache lifetimes
The lifetime — the time for which you trust a copy — is the proxy for the question: how much staleness can this value tolerate? It isn't a technical decision but a business one, and it should be made per data type, not globally.
A lifetime that is too long shows stale data and undermines trust in the system. A lifetime that is too short discards the copy before it has paid for itself and drives the hit rate down so far that the cache brings more overhead than benefit. There is no universal value between these poles; there is only the question of how much aging makes the specific value wrong enough to matter.
Two mechanisms soften hard expiry. The first is recomputing in the background before the copy expires: the value is refreshed while the old copy is still being served, so nobody waits on the computation. The second is tolerating slightly expired copies in case of failure: if the source is unreachable, you deliberately serve a stale copy instead of showing an error. Both shift the compromise in favor of availability — and must be chosen deliberately, because they further soften guaranteed freshness.
Trade-off. The lifetime trades freshness for hit rate and relief. Every extension increases relief and staleness at the same time.
Cost. Background recomputation adds complexity and shifts load to the time just before expiry; without jitter, this easily turns into a simultaneous rush on the source (see failure modes).
When we decide differently. For data that rarely does harm when wrong, we choose generous lifetimes and accept visible aging. For data close to money or law, we choose very short lifetimes or don't cache at all. For expensive but rarely changed aggregations, we combine a long lifetime with active invalidation on write — long enough to provide relief, but corrected immediately when the truth changes.
5. Layers of caching
Caching rarely happens in one place. Between the user and the data source sit several layers, and each can hold a copy. Knowing them matters, because a value can go stale in several layers at once and because control over invalidation diminishes from layer to layer.
Diagram: the closer the copy sits to the user, the greater the latency gain — and the less control you have over getting rid of it again.
| Layer | What it caches | Invalidation control | Blast radius of an error |
|---|---|---|---|
| Client / browser | responses per user | very low (someone else's device) | one user |
| Edge / CDN | shared, public responses | medium (purge, headers) | many users |
| Application (in-memory) | objects, computations | high (own process) | one node |
| Distributed cache | shared values across nodes | high, but coordinated | the entire system |
| Database / query | query results | medium (internal) | the source itself |
The client layer brings the biggest latency gain, because the request never even leaves the device — but a copy stored there is practically impossible to recall. That is why only things with a clearly bounded lifetime and no third-party identity belong on the client. The application layer is the easiest to control, but separate per node: in a system with multiple instances, each holds its own copy, and the same value can be of different ages on two nodes.
Trade-off. More layers mean more relief and more proximity to the user, but also more places where the same truth can age differently.
Cost. Every additional layer multiplies the effort of invalidation and complicates diagnosis: when an answer is wrong, you have to know which layer delivered it.
When we decide differently. We deliberately cache in as few layers as possible and choose the layer by purpose: public responses that are the same for everyone at the edge; expensive internal computations in the application. We only combine several layers for the same value if each has its own named purpose — never "just to be safe."
6. Distributed caches
As soon as a system consists of multiple nodes, a process-local cache is no longer enough, because each node holds its own truth. A distributed cache — storage shared by all nodes — solves this problem and creates a new one: the cache itself becomes a component that can fail, slow down and become overloaded. It is then no longer just an accelerator, but a part of the system on which availability depends.
That shifts the question. With a local cache it was: how do I keep my copy current? With a distributed cache, another question is added: what happens when the cache doesn't respond? A system that no longer works without its cache has merely swapped the database for another, often less durable dependency. The distributed cache must therefore be treated like any other network dependency: with timeouts, with defined behavior on failure and with the assumption that it will occasionally be gone.
A second peculiarity is consistency between cache and source under concurrency. When two nodes change the same value at the same time and update the cache, the order determines the result — and order across the network isn't guaranteed. The pattern "read, modify, write to the cache" is a race condition without additional safeguards. The more reliable basic form is to remove the copy on a change rather than overwrite it, so that the next reader fetches it fresh from the source.
Trade-off. A distributed cache trades inconsistency between nodes for a new, shared dependency and its risk of failure.
Cost. You are now operating an additional stateful system: it needs capacity planning, an eviction strategy, monitoring and a plan for when it is full or gone.
When we decide differently. As long as a system runs on one node or the nodes can tolerate small, acceptable deviations, we stick with the local cache — it is simpler and has no network dependency. We introduce the distributed cache when the deviation between nodes causes business problems or when the volume to be cached becomes too large for any single process.
7. Failure modes
A cache doesn't just change behavior in the normal case; it creates its own failure patterns that wouldn't exist without it. They typically occur under load — that is, precisely when the cache is supposed to help.
The best known is the simultaneous rush after expiry, often called a "stampede": a hot value expires, and all waiting requests hit the source at the same moment because none has written the new copy yet. The cache that was supposed to protect the source amplifies the load instead of damping it. Countermeasures are a lock, so that only one request recomputes while the others wait or receive the old copy, and jittering expiry times so that many values don't expire at once.
The second is the request for something that doesn't exist at all: if a value is missing at the source, the cache never finds it, and every request falls through to the source. If this state is triggered from outside, it can lead to overload. The countermeasure is to cache the absence as well — a short-lived "doesn't exist" entry — with the caveat that such an entry must not mask the value being created later.
The third is a total cache outage: if a shared cache layer fails, the full load suddenly hits the source, which was never designed for that load. A system that can't survive without its cache thus has a single point of failure. The countermeasure isn't in the cache but in how the source is designed: it must withstand a sudden load spike, at least in a throttled form.
Trade-off. The safeguards against these failures — locks, jitter, negative caching — increase complexity and can introduce new edge cases.
Cost. These failures usually only show up under real load and are hard to reproduce. The price of ignoring them is an outage at precisely the moment of highest demand.
When we decide differently. For rarely requested or non-critical values, we skip the elaborate safeguards and accept a brief extra load during the occasional rush. For hot, expensive values we build in locking and jitter from the start, because there the stampede isn't a question of if, but when.
8. Observability
A cache you can't measure is a guess. Without metrics, you know neither whether it helps nor whether it is currently serving wrong data. Observability is therefore not an add-on, but the prerequisite for being able to take responsibility for a cache at all.
The first metric is the hit rate — the share of requests served from the cache. It tells you whether the cache is fulfilling its purpose. A persistently low hit rate means you're bearing the cost of the copy without reaping its benefit, and it is a signal to reconsider the lifetime, the key design or the decision to cache at all. The second is the eviction rate — how often entries are removed due to lack of space — because high eviction silently undermines the hit rate. The third is the latency of the cache itself: a distributed cache that responds slowly can be more expensive than the source it is supposed to replace.
Harder but more important is observing correctness. Hit rates say nothing about whether the hits were correct. That is why a serious cache includes the ability to make the age and origin of a response visible — which layer it came from and how old the copy was — so that when an answer is wrong, you find the responsible layer instead of guessing.
Trade-off. Measurement and tagging cost some compute time and memory and add instrumentation to the hot path.
Cost. Without this instrumentation, cache errors are among the hardest problems to diagnose, because they manifest as sporadically wrong data without an error message.
When we decide differently. For a small, local cache with a clearly bounded effect, simple hit counters are enough. For every shared or distributed cache, we consider hit rate, eviction, latency and the origin of the response mandatory — in the spirit of observability as an architectural principle, not as a later add-on.
9. Common mistakes
The recurring patterns where caching fails — almost all are variants of the same mistake, treating the cache as free:
- Caching without measuring first — relieving the symptom and preserving the actual cause at the source.
- Caching permissions, tenant-bound or money-related data, thereby trading correctness for latency where that isn't allowed.
- Leaving the identity out of the cache key, so that one user receives another user's copy.
- Thinking only about writing the copy, not about invalidating it — every overlooked write path leaves a wrong copy behind.
- Choosing one global lifetime for everything instead of aligning it per data type with the tolerable staleness.
- Not jittering expiry times, so that many values expire simultaneously and hit the source at the same moment.
- Treating the distributed cache like a safe source — without a timeout, without defined behavior when it fails.
- Overwriting the copy on changes instead of removing it, thereby building in a race condition under concurrency.
- Not caching the absence of a value, letting repeated misses fall through to the source.
- Not measuring the cache, and therefore noticing neither benefit nor harm until users report wrong data.
10. Decision checklist
Clarify the following, in order, before introducing a cache:
- Measured? Is it established what is slow and why — and is it ruled out that the cause would be better fixed at the source (index, query, synchronous call)?
- May this value age? Is a slightly stale answer harmless here — and is the value not in the realm of permissions, tenant isolation or money?
- Key complete? Does the cache key contain everything that distinguishes the response, especially the identity of the user or tenant?
- Invalidation decided? Is it clear whether invalidation happens via expiry, actively or event-driven — and, with active invalidation, are all write paths covered?
- Lifetime per data type? Is the duration derived from the tolerable staleness rather than set globally?
- Layer chosen deliberately? Does the copy sit in the layer that fits the purpose — and not in several "just to be safe"?
- Failure modes considered? Are stampede (lock, jitter), misses (negative caching) and cache outage (source survives the load spike) handled?
- Distributed cache as a dependency? Does the shared cache have a timeout and defined behavior when it fails?
- Observable? Are hit rate, eviction, latency and the origin/age of a response measured?
- Can it be removed? Can the cache be switched off if in doubt without the system grinding to a halt?
Anyone who can't answer these questions cleanly doesn't have a performance feature, but a second source of truth with no plan for maintaining it.
FAQ
Isn't caching the fastest way to speed up a slow system? It is the fastest way to hide slowness. Whether that's good depends on the cause. If it is sheer, repeated read load, the cache is the right answer. If it is a missing index or a synchronous call, the cache preserves the problem and makes it harder to find later. That is why measurement comes before the cache.
Why is cache invalidation so hard? Because it requires knowing in one place that something has changed in another. Every path through which the truth can change must invalidate the copy — and a single overlooked write path is enough for a wrong copy to remain. The bug then only shows up sporadically, which makes it particularly stubborn.
Should we set a global lifetime? No. The lifetime is a business statement about how much staleness a particular value can tolerate, and that differs from data type to data type. A global value is either too long for the sensitive data or too short for the non-critical data — usually both at once.
Do we overwrite or delete the copy on a change? When in doubt, delete. Overwriting under concurrency is a race: two simultaneous changes can leave the cache in an order that contradicts the source. Removing the copy forces the next reader to fetch it fresh and is therefore more robust — at the price of one additional miss.
What happens if our distributed cache fails? That has to be answered before the outage. A cache is a network dependency: it needs a timeout and defined behavior when it doesn't respond — usually: keep working by going straight to the source. A system that stops without its cache has swapped the database for a less durable dependency.
When is a distributed cache worth it instead of a local one? When the deviation between nodes causes business problems or when the volume to be cached becomes too large for any single process. As long as a local cache is enough, it is the better choice, because it is simpler and doesn't introduce a shared source of failure.
Further reading
- Software that still runs in ten years — why a second source of truth shapes changeability over a system's lifetime.
- Queue or synchronous processing and Introducing asynchronous processing — the related question of not doing work in the request path.
- Idempotency in distributed systems — why removing a copy is more robust than overwriting it is the same discipline of repeatability.
- Observability as an architectural principle — a cache you don't measure is a guess.
It is grounded in the Batunet Engineering Method: measure first, then decide; design for failure; build in small, reversible steps.
A cache is borrowed speed. You pay it back with consistency — the only question is whether deliberately and as planned, or by surprise and during an outage.
Referenced entities
Continue your engineering journey.
Related concepts, decisions, playbooks and perspectives — as one connected path, not a list of links.
Engineering decisions
Playbooks
Perspectives
A concrete project in this field?
Reference Guides show how we think. For your system, talk to our management — technical, no sales pitch.
