DeploymentsReplicas and scaling

Replicas and scaling

Scale a deploy out by running more replicas, or up by giving each replica more CPU and memory.

Two ways to scale#

Scale outScale up
What changesThe number of replicasCPU and memory per replica
WhereEdit → ReplicasAdjust resources
Good forMore concurrent traffic, and staying available while one replica restartsA single replica running out of CPU or memory
Shared billingEach replica is billed, so 2 replicas cost 2×Billed on the larger request, per replica

Set the replica count#

  1. Edit the deploy

    In Deployments, choose Edit on the deploy.

  2. Change Replicas

    Set Replicas to the number of copies you want running. The minimum is 1.

  3. Save

    Choose Save deploy. The Replicas column updates, and the Pods page lists one pod per replica once they're running.

All replicas of a deploy run in the deploy's one location, and every replica gets the full CPU and memory of the build. Replicas share traffic from the deploy's routes.

Per-replica and total resources#

The Resources column shows CPU and memory per replica, with a /replica suffix when there's more than one. The Resources drawer shows both: Per replica, and Total across the replicas.

Example
Per replica          0.5 CPU · 512 MB
Total across 3       1.5 CPU · 1.5 GB

On shared compute you're billed per second on the total. See Usage and billing.

Resize each replica#

Choose Adjust resources on the deploy to change CPU and memory, or pick a Quick preset (Nano, Small, Medium or Large). Shared deploys can request 0.125–4 CPU and 256 MB–8 GB per replica, with 1–4 GB of memory per CPU. Dedicated deploys are capped to what their machine type can fit. Applying the change creates a new build and redeploys, and the workload may restart. See CPU and memory.

Autoscaling#

A deploy can have autoscaling (horizontal pod autoscaling) enabled. When it is, the Replicas column shows the range, for example 2 (HPA 2–6), and the Resources drawer shows Autoscaling with the replica range and the total resources across it. Saving the deploy form keeps existing autoscaling settings unchanged.

Scaling on isolated compute#

On dedicated tenancy you pay for nodes rather than replicas. When your workloads request more compute than a node has, another node is leased automatically and released when it's no longer needed, and each node is billed per second while it's leased. See Isolated compute.