Back to Blog
2026-07-12 · 8 min read

Incident_Architecture_Breakdown

Scale to Zero Without Losing Work: The Queue, Wake-Up, and Readiness Path

How I added scale-to-zero to eligible worker workloads while keeping accepted jobs durable and making cold-start recovery explicit.

Scale to ZeroKubernetesAWS ECSQueuesReliabilityFinOps

1. Hook and Stakes

Keeping asynchronous workers running while no jobs are queued wastes capacity, but setting replicas to zero can lose work or create a broken first request if the wake-up path is incomplete.

A credible scale-to-zero design must preserve accepted work, distinguish eligible workloads, recover from zero capacity, and expose the cold-start tradeoff instead of hiding it.

2. Architecture Diagram

The control plane and durable queue stay available while eligible execution workers scale independently between zero and bounded capacity.

mermaid
graph LR
  Client[Client]-->API[Always-Available Control API]
  API-->Queue[Durable Queue + DLQ]
  Queue-->Metric[Backlog / Oldest-Job Signal]
  Metric-->Policy[Idle Window + Scaling Policy]
  Policy-->Workers[Eligible Workers: 0 to N]
  Workers-->Ready[Readiness Gate]
  Ready-->Execute[Claim and Execute Job]
  Execute-->Results[(Result Store)]
  • Always-available request intake separated from elastic execution capacity
  • Durable queue and DLQ state that survive zero active workers
  • Opt-in workload eligibility plus configurable idle window and replica floor
  • Backlog-triggered 0-to-N wake-up with cooldown and hysteresis
  • Readiness gate before a cold worker can claim queued work

3. Stress Test and Breaking Point

Setup: I modeled the full lifecycle from an empty queue and zero eligible workers through a new job arrival, worker startup, readiness, execution, and return to idle.

Failure Signal: A replica count of zero is unsafe if scaling depends only on worker CPU, because no running worker exists to emit the signal needed to wake itself.

  • Queue backlog and oldest-job age provide external wake-up signals even when execution capacity is zero.
  • Accepted jobs remain queued while replacement capacity starts instead of being coupled to a synchronous request timeout.
  • Readiness gating keeps warming workers from claiming jobs before their runtime dependencies are available.
  • Per-workload eligibility prevents stateful or latency-sensitive services from inheriting a zero-replica minimum.

4. Bottleneck Root Cause and Resolution

Root Cause: CPU-based autoscaling cannot recover a service from zero because there are no active tasks or pods producing CPU metrics.

Resolution: I moved the wake-up decision to durable external signals: queue backlog, oldest-job age, workload eligibility, and an idle timer. The control loop raises desired capacity, waits for readiness, and only then allows workers to consume jobs.

  • Scale-to-zero removes idle worker allocation but adds cold-start latency to the first queued job after inactivity.
  • Short idle windows save capacity sooner but can create churn, so cooldown and hysteresis must be tuned together.
  • Not every service is eligible; stateful and latency-sensitive workloads need an explicit minimum replica floor.

AWS Alternatives Considered

  • Keep one warm worker for latency-sensitive queues and use scale-to-zero only for batch or delay-tolerant workloads.
  • Use Kubernetes event-driven autoscaling for queue-backed workloads when cluster-level control is preferred.
  • Use managed serverless execution when its runtime, isolation, and startup constraints match the workload.

5. Business Impact

  • Reduces idle execution allocation for bursty, delay-tolerant workloads.
  • Preserves reliability by keeping intake and queue state independent from worker count.
  • Makes the cost-versus-latency decision visible enough for reviewers to challenge in a system-design interview.

References and Live Evidence