Autoscaling self-hosted runners Jump to heading

Self-hosted runners usually start as a few long-lived machines registered by hand. That works until load varies: at 10:00 every runner is busy and jobs queue for fifteen minutes; at 03:00 the machines sit idle and cost money. Long-lived runners also accumulate state between jobs β€” leftover files, background processes, poisoned caches β€” which makes builds flaky and, for any repository that accepts outside contributions, unsafe. Autoscaling ephemeral runners fixes both: a controller watches the job queue, creates a fresh runner per job from a clean image, and destroys it afterwards. The work is in keeping that fast β€” clean machines start cold β€” and in setting limits so a burst of jobs does not become a burst of spend. This page covers the setup, within CI caching and runner performance.

When to use this approach Jump to heading

  • You use self-hosted runners for cost, hardware, network access or licensing reasons.
  • Job queue times spike during the working day while runners sit idle at night.
  • Builds are flaky because state leaks between jobs on long-lived runners.
  • Provenance or security requirements call for ephemeral runners, as in attesting build artefacts from a self-hosted runner.

Step 1 β€” Make runners ephemeral Jump to heading

An ephemeral runner registers, takes exactly one job, and deregisters. Everything else in this page builds on that property, because it is what lets a controller treat runners as disposable.

# Runner start-up script inside the image: register for one job, run it, exit
./config.sh --url "https://github.com/$ORG" --token "$REG_TOKEN" \
  --labels "linux,x64,build" --ephemeral --unattended --disableupdate
./run.sh
# when run.sh exits, the orchestrator destroys this VM or pod
Long-lived runners against ephemeral runnersLong-lived runners keep files, processes and caches between jobs, so builds can affect each other and capacity is fixed. Ephemeral runners start clean for every job and are destroyed afterwards, so jobs are isolated and capacity can follow demand.Long-livedEphemeralstate between jobsleaksnonecapacityfixedfollows the queuestart-up timeinstantimage bootsafe for untrusted PRsnowith isolationephemeral runners trade start-up time for isolation β€” step 3 wins most of it back

Step 2 β€” Let a controller scale on queued jobs Jump to heading

A controller watches for queued jobs that match your runner labels and creates runners to serve them. On Kubernetes, Actions Runner Controller does this with runner scale sets; on VMs, cloud autoscaling groups driven by webhook events or the API serve the same purpose.

# Actions Runner Controller β€” a runner scale set (Helm values excerpt)
githubConfigUrl: https://github.com/acme
githubConfigSecret: arc-github-app
minRunners: 2          # a small warm pool for the working day
maxRunners: 40         # hard ceiling on cost and on load to shared services
template:
  spec:
    containers:
      - name: runner
        image: registry.example.com/runners/build:2026.10
        resources: { requests: { cpu: "4", memory: 8Gi } }
# Workflows target the scale set by its name
jobs:
  build:
    runs-on: build-linux

Step 3 β€” Win back start-up time with pre-built images Jump to heading

Ephemeral runners are only fast if starting one is fast. Bake everything a job needs that does not change per job β€” toolchains, language runtimes, common CLI tools, a Git mirror of large repositories β€” into the runner image, rebuilt on a schedule.

FROM ghcr.io/actions/actions-runner:2.319.1
RUN sudo apt-get update && sudo apt-get install -y --no-install-recommends build-essential jq git-lfs \
 && sudo rm -rf /var/lib/apt/lists/*
# A read-only reference mirror speeds up clones of the monorepo
RUN git clone --mirror https://github.com/acme/monorepo.git /opt/mirrors/monorepo.git
# In the job: clone using the baked mirror as a reference, then fetch only what is new
git clone --reference-if-able /opt/mirrors/monorepo.git "$REPO_URL" work
git -C work checkout --detach "$COMMIT_SHA"

A reference mirror baked into the image gives every job most of the repository’s objects without network transfer; the technique is covered in reusing a Git mirror on self-hosted runners.

From queued job to finished runnerA job is queued with matching labels. The controller creates a runner from a pre-built image that already contains toolchains and a repository mirror. The runner registers, takes the single job, runs it with warm local data, and is destroyed when the job ends.Job queuedruns-on labelControllerscale upPre-built imagetools + mirrorOne job--ephemeralDestroyedno state kepteverything that does not vary per job belongs in the image, not in the job

Step 4 β€” Share caches without sharing state Jump to heading

Ephemeral runners cannot keep a local cache, but they can read and write a remote one: the CI provider’s cache service, an S3-compatible bucket, or a BuildKit registry cache. Keep write access scoped so jobs from pull requests cannot overwrite caches used by the default branch.

      - uses: actions/cache@v4
        with:
          path: ~/.cache/pip
          key: pip-${{ runner.os }}-${{ hashFiles('**/requirements*.txt') }}

Container build caches follow the same pattern, described in caching Docker layers in CI.

Step 5 β€” Set limits and watch the cost Jump to heading

Autoscaling without limits turns a runaway workflow β€” a matrix accidentally expanded to hundreds of jobs β€” into a large bill and a flood of load on shared services such as package registries. Set a maximum per scale set, a minimum warm pool sized for normal mornings, and scale-down delays that avoid thrashing.

# Queue time and runner count over the day, from the CI API
gh api "orgs/$ORG/actions/runners" --jq '.runners | group_by(.status) | map({(.[0].status): length}) | add'
gh run list --limit 200 --json createdAt,startedAt \
  --jq '[.[] | ((.startedAt|fromdateiso8601) - (.createdAt|fromdateiso8601))] | (add/length) | "avg queue wait: \(. | floor) s"'
Queue wait before and after autoscalingAn illustrative team on six fixed runners waited many minutes for a runner at peak and paid for idle machines overnight. With an autoscaled scale set β€” a warm pool of two and a ceiling of forty β€” peak queue wait dropped to under a minute.average queue wait in minutes (illustrative)fixed, 10:0014 minfixed, 03:000autoscaled, 10:000.7 minautoscaled, 03:000.4 minthe warm pool covers the first jobs of the morning while new runners boot

Validation checklist Jump to heading

Frequently Asked Questions Jump to heading

How small can the warm pool be? Jump to heading

Size it for the burst at the start of the working day, while new runners boot. Two to five runners is common for medium teams; measure queue wait at 09:00 and adjust.

Do ephemeral runners make builds slower? Jump to heading

The first minute of each job can be slower than on a warm long-lived machine. Pre-built images and remote caches recover most of that, and the reliability gain from clean state usually saves more time than it costs.

Can one scale set serve every repository? Jump to heading

It can, but separating scale sets by trust level β€” internal repositories, public repositories accepting outside contributions, release builds β€” keeps a compromise in one from reaching the others.