Zero-downtime deployment architecture with health checks, rolling updates, and backwards compatibility
DevOps 8 min read

What Zero-Downtime Deployment Actually Requires

Zero-downtime deployment means releasing new code without interrupting active requests or making the service unavailable. It requires health checks, rolling deployments, backwards-compatible changes, connection draining, and monitoring. Zero-downtime is an architectural property, not merely a deployment tool feature.

Health checks

Health checks tell load balancers and orchestrators when instances are ready to receive traffic.

Readiness checks

Readiness checks verify the application can handle requests: database connections established, dependencies available, caches warmed. New instances should not receive traffic until ready.

Liveness checks

Liveness checks detect crashed or hung processes. Failed liveness checks trigger instance replacement.

Implementing health endpoints

GET /health/ready
200 OK if ready, 503 Service Unavailable if not

GET /health/live  
200 OK if alive, 503 if crashed

Rolling deployments

Rolling deployments update instances gradually rather than all at once.

Rolling deployment process

  1. Deploy new version to one instance
  2. Wait for health checks to pass
  3. Route traffic to new instance
  4. Repeat for remaining instances
  5. Keep old version running until rollout completes

Blue-green deployments

Blue-green maintains two complete environments. Deploy to inactive environment, test, then switch traffic. Enables instant rollback but requires double infrastructure.

Backwards-compatible changes

During rolling deployments, old and new code run simultaneously. Changes must be compatible with both versions.

Database migrations

Breaking database changes require multi-step migrations:

  1. Add new column (old code ignores it)
  2. Deploy code that writes to both old and new columns
  3. Backfill data
  4. Deploy code that reads from new column
  5. Remove old column

API changes

API changes must be backwards-compatible or versioned. Removing fields or changing behaviour breaks existing clients.

Connection draining

Shutting down instances immediately terminates active requests. Connection draining allows in-flight requests to complete.

Graceful shutdown

  1. Stop accepting new requests
  2. Allow existing requests to complete (with timeout)
  3. Close database connections
  4. Shut down

Monitoring during deployment

Zero-downtime requires observing deployment health in real-time.

Key metrics

  • Error rate (should not spike)
  • Latency (should remain stable)
  • Request success rate
  • Health check status

Automated rollback

If error rates increase during deployment, automatically roll back to previous version.

Zero-downtime is architectural

Zero-downtime deployment requires health checks, rolling updates, backwards-compatible changes, connection draining, and monitoring. It is not a single tool feature but a set of architectural decisions. For broader deployment practices, see Observable Deployments.


Published by the DSSS Engineering Team. For corrections or topic requests, use the contact page.