Backend
Lesson 6 of 8About 4 min readSuggest an edit

Scaling a service

Scaling means handling more load without getting slower or less reliable. Adding servers only helps if servers are the bottleneck, so find where time is spent first, then make the cheapest change that removes the constraint.

Find the bottleneck first

A request’s latency is the sum of its parts: application CPU, database queries, calls to other services, and time spent waiting for a free connection or thread. Measure before you change anything.

  • Look at latency percentiles (p50, p95, p99), not averages, which hide the slow tail.
  • Trace a slow request to see which step dominates.
  • Check saturation: CPU, memory, database connections, disk I/O, queue depth.

Often the fix is not hardware at all, but a missing index, an N+1 query or a loop of remote calls.

Vertical and horizontal scaling

Aspect Vertical Horizontal
What a bigger machine more machines
Effort small, often a config change needs a stateless design
Limit the largest machine you can buy coordination and shared resources
Failure one machine is still one failure losing one instance is tolerable

Vertical scaling is underrated: it is simple and buys time. Horizontal scaling is how you grow past one machine and survive the loss of one.

Stateless services

To add instances freely, each one must be stateless: it keeps nothing between requests that another instance would need. Sessions, uploads and caches that must be shared go into external stores such as a database, Redis or object storage. Then any instance can serve any request, and you can add or remove instances without users noticing.

Load balancing

A load balancer spreads requests across healthy instances. Round-robin or least-connections are common strategies. It runs health checks and stops sending traffic to instances that fail them. Keep the health check cheap and honest: it should fail when the instance can’t serve, but not because a single optional dependency is slow, or one slow service will pull your whole fleet out of rotation.

Avoid sticky sessions where you can; they make load and deploys uneven.

Connection pooling

Opening a database connection is expensive, and databases limit how many they accept. A connection pool keeps a set of open connections that requests borrow and return.

database:
  pool_min: 2
  pool_max: 10
  acquire_timeout_ms: 2000

Watch the arithmetic when you scale out: 40 instances with a pool of 10 each is 400 connections, which may exceed what the database allows. A bigger pool is not always faster either; beyond a point, more concurrent queries just contend for the same CPU and disk. A pooler in front of the database, such as PgBouncer for PostgreSQL, lets many application connections share fewer database ones.

Read replicas and replication lag

When the database is the bottleneck and most traffic is reads, add read replicas: copies that receive changes from the primary and serve read queries. Writes still go to the primary.

Replication is usually asynchronous, so replicas trail the primary by some amount of replication lag. A user who saves a profile and immediately reloads it from a replica may see the old version. Route reads that must see the user’s own writes to the primary, at least for a short period after a write.

Sharding basics

When one primary can’t handle the writes or hold the data, sharding splits rows across several databases by a shard key, such as tenant_id. Each shard owns a subset of the data.

Choose the key so that most queries touch one shard and load spreads evenly. Queries across shards, joins across shards and transactions across shards become hard or impossible, and moving data between shards later is a serious project. Treat sharding as a late step, after indexing, caching, replicas and vertical scaling.

How to decide

  1. Measure and find the saturated resource.
  2. Fix inefficiency: queries, indexes, caching.
  3. Scale vertically if it is cheap and quick.
  4. Make the service stateless and scale it horizontally.
  5. Offload reads to replicas, handling lag.
  6. Shard only when writes or data size outgrow a single primary.

Next: Timeouts, retries and circuit breakers

Timeouts on every call, safe retries with backoff and jitter, circuit breakers, bulkheads and fallbacks.