Deployment strategies
Every deployment carries some risk, because production traffic and data find problems that tests miss. A deployment strategy decides how much of that risk users are exposed to at once, and how fast you can back out.
Recreate
Stop all old instances, then start the new ones. It is simple and guarantees two versions never run together, but users see downtime between the two steps. It suits internal tools and cases where two versions cannot coexist.
Rolling
Replace instances a few at a time, waiting for each new one to pass its health check before continuing. It is the default in Kubernetes and most platforms, and causes no downtime. The cost is that old and new versions serve traffic side by side for the length of the rollout, and rolling back is another rollout, not an instant switch.
Blue-green
Run two complete environments. Blue serves production; you deploy the new version to green, test it there, then switch the router or load balancer to send all traffic to green. Rolling back means switching back, which takes seconds as long as blue is still running. The price is double capacity during the switch, and both environments usually share one database, so the schema must work for both versions.
Canary
Send a small share of traffic, say 1% or 5%, to the new version and compare its error rate and latency with the current version. If they look healthy, increase the share in steps until it reaches 100%. A canary limits a bad release to a fraction of users. It needs weighted traffic splitting and enough traffic for the comparison to mean something.
| Strategy | Downtime | Extra capacity | Rollback speed |
|---|---|---|---|
| Recreate | Yes | None | Slow (redeploy) |
| Rolling | No | Small | Medium (roll back) |
| Blue-green | No | Double, briefly | Fast (switch) |
| Canary | No | Small | Fast (shift traffic) |
Feature flags: deploy is not release
A feature flag lets code reach production switched off. Deploying puts the code on servers; releasing turns the behaviour on for users. Separating them lets you release to staff, then a percentage of users, and turn a feature off in seconds without deploying.
Flags have a cost: every flag doubles the paths through the code it guards. Give each release flag an owner and a removal date, and delete it once the feature is fully rolled out.
Database migrations with expand and contract
During any strategy except recreate, old and new code run against the same database. A migration that renames or drops a column breaks whichever version does not expect it. The expand and contract pattern splits a breaking change into safe steps:
1. Expand: add the new column (nullable, or with a default)
2. Deploy: code writes to both columns, reads the old one
3. Backfill: copy existing data into the new column
4. Deploy: code reads the new column
5. Contract: after a full release, stop writing and drop the old
Each step works with the code before and after it, so every deploy can be rolled back. Watch out for migrations that lock large tables; on big tables, add indexes and backfill data in ways that don’t block writes.
Automated rollback on metrics
Define in advance what “unhealthy” means for a rollout, for example an error rate or p99 latency clearly worse than the baseline, and let the pipeline or a progressive delivery tool (Argo Rollouts, Flagger, Spinnaker and similar) check it at each step. If a check fails, traffic goes back to the old version automatically and someone is notified. Rollback must be safe for this to work, which is why migrations and flags matter.
How to decide
- Internal tool, downtime acceptable: recreate.
- Most stateless services: rolling, with good readiness checks.
- Need an instant switch-back and can afford the capacity: blue-green.
- High traffic, high cost of a bad release: canary with automated analysis.
- Whatever you choose, use feature flags for risky behaviour and expand and contract for schema changes.