Caching
A cache keeps a copy of a result so the next request can skip the work that produced it. Done well, it cuts latency and database load dramatically. Done carelessly, it serves stale prices or leaks one user’s data to another.
Where caches live
Caches sit at every layer between the user and the data.
| Layer | Examples | Good for |
|---|---|---|
| HTTP and CDN | browser cache, CDN edge | public, shared responses |
| Application | in-process map, Redis, Memcached | computed results, hot rows |
| Database | buffer pool, materialised views | repeated reads of the same pages |
HTTP caching is controlled by headers such as Cache-Control and ETag (see the HTTP fundamentals lesson). An in-process cache is the fastest but each instance has its own copy, so instances can disagree. A shared cache such as Redis gives every instance the same view at the cost of a network hop.
Cache-aside and write-through
The most common pattern is cache-aside: the application checks the cache, and on a miss reads the database and fills the cache.
def get_product(product_id):
key = f"product:{product_id}"
cached = redis.get(key)
if cached is not None:
return json.loads(cached)
product = db.fetch_product(product_id)
redis.set(key, json.dumps(product), ex=300)
return product
On writes, update the database and then delete the cache key, so the next read loads fresh data. Deleting is safer than writing the new value into the cache, because two concurrent writers can otherwise leave the older value in place. A smaller race remains (a slow reader can put an old value back just after the delete), which is one reason every entry still needs a TTL.
With write-through, every write goes to the cache and the database together, so the cache is always populated. Reads are fast from the start, but you pay for caching data that may never be read, and you still need a plan for when one of the two writes fails.
TTLs and invalidation
Invalidation is the hard part. You have two tools:
- A TTL (time to live) bounds how stale an entry can get. It is your safety net: even if an invalidation is missed, the entry expires.
- Explicit invalidation deletes keys when the underlying data changes. It keeps data fresh but requires you to know every key a write affects.
Use both. Pick the TTL by asking how stale the data can be before a user notices or is harmed. A product description can be minutes old; a stock count at checkout should not be cached at all.
Cache stampedes
When a popular key expires, every request that misses at that moment goes to the database at once. This is a cache stampede, and it can take down the database the cache was protecting. Several defences combine well:
- Request coalescing: only one caller recomputes a missing key while the others wait for its result. In a single process this is a per-key lock or shared promise; across instances, a short-lived lock in the shared cache.
- Jitter: add a random amount to each TTL so keys written together don’t expire together.
- Stale-while-revalidate: keep serving the expired value while one background refresh fetches a new one. CDNs support this through the
stale-while-revalidatedirective ofCache-Control.
Cache-Control: max-age=60, stale-while-revalidate=30
What not to cache
The most damaging caching bug is serving one user’s data to another. It happens when a response depends on the user but the cache key does not include them, for example a CDN caching /account because nobody set the right header.
- Mark per-user responses
Cache-Control: privateorno-store, so shared caches never store them. - Include every input that changes the output in the cache key: user or tenant ID, locale, permissions.
- Don’t cache secrets, tokens or payment details in a shared cache.
Checklist before adding a cache
- Measure first: confirm the slow path is actually read often with the same inputs.
- Write down the key format and every input it depends on.
- Choose a TTL from the staleness you can tolerate, and add jitter.
- Decide how writes invalidate entries, and delete rather than overwrite.
- Protect hot keys from stampedes.
- Make sure the service still works, only slower, if the cache is empty or down.