Backend
Lesson 4 of 8About 4 min readSuggest an edit

Caching

A cache keeps a copy of a result so the next request can skip the work that produced it. Done well, it cuts latency and database load dramatically. Done carelessly, it serves stale prices or leaks one user’s data to another.

Where caches live

Caches sit at every layer between the user and the data.

Layer Examples Good for
HTTP and CDN browser cache, CDN edge public, shared responses
Application in-process map, Redis, Memcached computed results, hot rows
Database buffer pool, materialised views repeated reads of the same pages

HTTP caching is controlled by headers such as Cache-Control and ETag (see the HTTP fundamentals lesson). An in-process cache is the fastest but each instance has its own copy, so instances can disagree. A shared cache such as Redis gives every instance the same view at the cost of a network hop.

Cache-aside and write-through

The most common pattern is cache-aside: the application checks the cache, and on a miss reads the database and fills the cache.

def get_product(product_id):
    key = f"product:{product_id}"
    cached = redis.get(key)
    if cached is not None:
        return json.loads(cached)
    product = db.fetch_product(product_id)
    redis.set(key, json.dumps(product), ex=300)
    return product

On writes, update the database and then delete the cache key, so the next read loads fresh data. Deleting is safer than writing the new value into the cache, because two concurrent writers can otherwise leave the older value in place. A smaller race remains (a slow reader can put an old value back just after the delete), which is one reason every entry still needs a TTL.

With write-through, every write goes to the cache and the database together, so the cache is always populated. Reads are fast from the start, but you pay for caching data that may never be read, and you still need a plan for when one of the two writes fails.

TTLs and invalidation

Invalidation is the hard part. You have two tools:

  • A TTL (time to live) bounds how stale an entry can get. It is your safety net: even if an invalidation is missed, the entry expires.
  • Explicit invalidation deletes keys when the underlying data changes. It keeps data fresh but requires you to know every key a write affects.

Use both. Pick the TTL by asking how stale the data can be before a user notices or is harmed. A product description can be minutes old; a stock count at checkout should not be cached at all.

Cache stampedes

When a popular key expires, every request that misses at that moment goes to the database at once. This is a cache stampede, and it can take down the database the cache was protecting. Several defences combine well:

  • Request coalescing: only one caller recomputes a missing key while the others wait for its result. In a single process this is a per-key lock or shared promise; across instances, a short-lived lock in the shared cache.
  • Jitter: add a random amount to each TTL so keys written together don’t expire together.
  • Stale-while-revalidate: keep serving the expired value while one background refresh fetches a new one. CDNs support this through the stale-while-revalidate directive of Cache-Control.
Cache-Control: max-age=60, stale-while-revalidate=30

What not to cache

The most damaging caching bug is serving one user’s data to another. It happens when a response depends on the user but the cache key does not include them, for example a CDN caching /account because nobody set the right header.

  • Mark per-user responses Cache-Control: private or no-store, so shared caches never store them.
  • Include every input that changes the output in the cache key: user or tenant ID, locale, permissions.
  • Don’t cache secrets, tokens or payment details in a shared cache.

Checklist before adding a cache

  • Measure first: confirm the slow path is actually read often with the same inputs.
  • Write down the key format and every input it depends on.
  • Choose a TTL from the staleness you can tolerate, and add jitter.
  • Decide how writes invalidate entries, and delete rather than overwrite.
  • Protect hot keys from stampedes.
  • Make sure the service still works, only slower, if the cache is empty or down.

Next: Queues and background jobs

Moving work out of the request, delivery guarantees, retries, dead-letter queues and the outbox pattern.