Networking for DevOps
Most “the site is down” incidents turn out to be about names, certificates or paths between machines rather than application code. You need to be able to follow a request from a browser to your process and know where it can fail.
DNS resolution and TTLs
When a client looks up api.example.com, it asks a recursive resolver (usually run by the ISP, the cloud provider or a public service). The resolver asks the root servers, then the servers for .com, then the authoritative servers for example.com, and caches the answer. Common record types are A (IPv4), AAAA (IPv6) and CNAME (an alias to another name).
Every record has a TTL, the number of seconds resolvers may cache it. A long TTL reduces lookups; a short one lets changes take effect quickly. Before a planned migration, lower the TTL a day or so ahead, wait for the old TTL to expire, make the change, then raise it again. Some clients cache longer than they should, so keep the old target running for a while.
TCP and UDP
TCP gives an ordered, reliable byte stream over a connection set up with a handshake. HTTP/1.1 and HTTP/2, databases and SSH use it. UDP sends independent datagrams with no delivery guarantee and no connection setup, which suits DNS queries, video, metrics and QUIC (the transport under HTTP/3).
TLS and certificates
TLS encrypts the connection and proves the server’s identity. The server presents its certificate plus intermediate certificates; the client builds a chain from them up to a root certificate it already trusts. A common mistake is serving only the leaf certificate: browsers may cope, while other clients fail with a verification error.
Certificates expire, and an expired certificate is a full outage that is entirely predictable. Automate renewal with ACME (Let’s Encrypt with certbot, cert-manager in Kubernetes) or use your cloud provider’s managed certificates, and still monitor expiry dates in case automation silently stops.
Load balancers and reverse proxies
A load balancer spreads traffic across backends that pass health checks.
| Aspect | Layer 4 | Layer 7 |
|---|---|---|
| Sees | IPs, ports, TCP/UDP | HTTP method, path, headers |
| Can route by | Address and port | Host, path, header, cookie |
| Typical use | Raw TCP, very high throughput | Web apps, APIs, TLS termination |
A reverse proxy (nginx, HAProxy, Envoy, Caddy) sits in front of your services and handles TLS termination, routing, compression, rate limiting and timeouts. Behind a proxy, your app sees the proxy’s IP as the client, so read the real client address from X-Forwarded-For, and trust that header only when it was set by your own proxy.
Private networks and firewalls
Put databases, caches and internal services in private subnets with no public IP addresses, and let only load balancers face the internet.
Cloud security groups are stateful firewalls attached to instances or interfaces: if an inbound connection is allowed, its replies are allowed automatically. Allow only what is needed, and where possible allow traffic from another security group rather than an IP range. A connection that hangs until it times out often means a firewall is silently dropping packets; an immediate “connection refused” usually means the host was reached but nothing is listening on that port.
Debugging with dig and curl
# What does DNS return, and with what TTL?
dig +noall +answer api.example.com
# Trace resolution from the root
dig +trace api.example.com
# Timing, TLS handshake and response headers
curl -sv -o /dev/null https://api.example.com/health
# Test one backend directly, bypassing DNS
curl --resolve api.example.com:443:10.0.1.12 \
https://api.example.com/health
curl -v shows the IP it connected to, the TLS version and the certificate, which settles most “DNS, TLS or app” questions. To see the full chain the server sends:
openssl s_client -connect api.example.com:443 -showcerts
A debugging order
- Does the name resolve to the address you expect?
- Can you open a TCP connection to that address and port?
- Does the TLS handshake succeed, with a valid chain and no expiry?
- Does the load balancer consider the backends healthy?
- Only then look at the application’s logs.