Kubernetes fundamentals
Kubernetes runs containers across a cluster of machines. You tell it the state you want, such as “three copies of this image, reachable on port 80”, and its controllers keep working to make the cluster match: restarting crashed containers, replacing lost nodes and rolling out new versions. That reconciliation loop is the main idea to hold on to.
The core objects
A pod is the smallest unit Kubernetes schedules: one or more containers that share a network address and can share volumes. Usually a pod holds one application container. Pods are disposable: a replacement pod gets a new IP address, so nothing should depend on a particular pod.
You rarely create pods directly. A deployment declares which image to run and how many replicas, and manages the pods for you.
A service gives a stable name and virtual IP to a changing set of pods, chosen by labels. Other workloads call http://web and the service forwards to whichever healthy pods currently match. External traffic arrives through an Ingress, the Gateway API or a LoadBalancer service.
apiVersion: apps/v1
kind: Deployment
metadata:
name: web
spec:
replicas: 3
selector:
matchLabels:
app: web
template:
metadata:
labels:
app: web
spec:
containers:
- name: web
image: registry.example.com/web:3f9c2a1
ports:
- containerPort: 3000
resources:
requests:
cpu: 250m
memory: 256Mi
limits:
memory: 512Mi
readinessProbe:
httpGet:
path: /ready
port: 3000
livenessProbe:
httpGet:
path: /healthz
port: 3000
---
apiVersion: v1
kind: Service
metadata:
name: web
spec:
selector:
app: web
ports:
- port: 80
targetPort: 3000
Configuration and secrets
A ConfigMap holds non-secret configuration, and a Secret holds credentials. Both can be exposed to containers as environment variables or mounted files, so the same image runs in every environment. Be aware that Secrets are only base64-encoded by default, not encrypted. Enable encryption at rest, restrict access with RBAC, and consider an external secrets manager.
Requests and limits
Requests are what the scheduler reserves for a container: a pod is placed only on a node with that much CPU and memory unreserved. Limits are the ceiling at runtime. The two behave differently when exceeded:
| Resource | Over the limit |
|---|---|
| CPU | The container is throttled and slows down |
| Memory | The container is killed (OOMKilled) and restarted |
Always set requests, or the scheduler packs pods blindly and they compete for resources. Set a memory limit based on observed usage. Many teams skip CPU limits because throttling causes hard-to-diagnose latency spikes; that is reasonable if requests are accurate.
Liveness and readiness probes
A readiness probe answers “should this pod receive traffic now?”. When it fails, the pod is removed from the service’s endpoints but left running. Use it for warm-up and for temporary overload.
A liveness probe answers “is this process stuck beyond recovery?”. When it fails, the container is restarted. Keep it simple and local. A liveness check that calls the database will restart every pod at once when the database is slow, turning a small problem into an outage. For slow-starting apps, add a startup probe so liveness checks don’t kill them before they are up.
Rolling updates
By default a deployment replaces pods gradually. maxSurge controls how many extra pods may be created, and maxUnavailable how many may be missing during the rollout. New pods receive traffic only once they are ready, which is why a real readiness probe matters. Your app must also handle SIGTERM and finish in-flight requests, as old pods are stopped.
kubectl rollout status deployment/web
kubectl rollout undo deployment/web
When Kubernetes is overkill
Kubernetes has real operating costs: upgrades, networking, security policies and people who understand them. For a handful of services, a managed container platform (Cloud Run, ECS on Fargate, App Runner, Fly.io and similar) or plain VMs with a deploy script are often the better choice.
How to decide
- Many services, several teams and a need for a shared platform: Kubernetes earns its keep.
- A few services and a small team: use a managed platform and revisit later.
- If you adopt it, use a managed cluster, set requests on every workload, and write probes that reflect real health.