Zero-downtime deployment means releasing new code without interrupting active requests or making the service unavailable. It requires health checks, rolling deployments, backwards-compatible changes, connection draining, and monitoring. Zero-downtime is an architectural property, not merely a deployment tool feature.
Health checks
Health checks tell load balancers and orchestrators when instances are ready to receive traffic.
Readiness checks
Readiness checks verify the application can handle requests: database connections established, dependencies available, caches warmed. New instances should not receive traffic until ready.
Liveness checks
Liveness checks detect crashed or hung processes. Failed liveness checks trigger instance replacement.
Implementing health endpoints
GET /health/ready
200 OK if ready, 503 Service Unavailable if not
GET /health/live
200 OK if alive, 503 if crashed
Rolling deployments
Rolling deployments update instances gradually rather than all at once.
Rolling deployment process
- Deploy new version to one instance
- Wait for health checks to pass
- Route traffic to new instance
- Repeat for remaining instances
- Keep old version running until rollout completes
Blue-green deployments
Blue-green maintains two complete environments. Deploy to inactive environment, test, then switch traffic. Enables instant rollback but requires double infrastructure.
Backwards-compatible changes
During rolling deployments, old and new code run simultaneously. Changes must be compatible with both versions.
Database migrations
Breaking database changes require multi-step migrations:
- Add new column (old code ignores it)
- Deploy code that writes to both old and new columns
- Backfill data
- Deploy code that reads from new column
- Remove old column
API changes
API changes must be backwards-compatible or versioned. Removing fields or changing behaviour breaks existing clients.
Connection draining
Shutting down instances immediately terminates active requests. Connection draining allows in-flight requests to complete.
Graceful shutdown
- Stop accepting new requests
- Allow existing requests to complete (with timeout)
- Close database connections
- Shut down
Monitoring during deployment
Zero-downtime requires observing deployment health in real-time.
Key metrics
- Error rate (should not spike)
- Latency (should remain stable)
- Request success rate
- Health check status
Automated rollback
If error rates increase during deployment, automatically roll back to previous version.
Zero-downtime is architectural
Zero-downtime deployment requires health checks, rolling updates, backwards-compatible changes, connection draining, and monitoring. It is not a single tool feature but a set of architectural decisions. For broader deployment practices, see Observable Deployments.
Published by the DSSS Engineering Team. For corrections or topic requests, use the contact page.