Autoscaling self-hosted runners Jump to heading
Self-hosted runners usually start as a few long-lived machines registered by hand. That works until load varies: at 10:00 every runner is busy and jobs queue for fifteen minutes; at 03:00 the machines sit idle and cost money. Long-lived runners also accumulate state between jobs β leftover files, background processes, poisoned caches β which makes builds flaky and, for any repository that accepts outside contributions, unsafe. Autoscaling ephemeral runners fixes both: a controller watches the job queue, creates a fresh runner per job from a clean image, and destroys it afterwards. The work is in keeping that fast β clean machines start cold β and in setting limits so a burst of jobs does not become a burst of spend. This page covers the setup, within CI caching and runner performance.
When to use this approach Jump to heading
- You use self-hosted runners for cost, hardware, network access or licensing reasons.
- Job queue times spike during the working day while runners sit idle at night.
- Builds are flaky because state leaks between jobs on long-lived runners.
- Provenance or security requirements call for ephemeral runners, as in attesting build artefacts from a self-hosted runner.
Step 1 β Make runners ephemeral Jump to heading
An ephemeral runner registers, takes exactly one job, and deregisters. Everything else in this page builds on that property, because it is what lets a controller treat runners as disposable.
# Runner start-up script inside the image: register for one job, run it, exit
./config.sh --url "https://github.com/$ORG" --token "$REG_TOKEN" \
--labels "linux,x64,build" --ephemeral --unattended --disableupdate
./run.sh
# when run.sh exits, the orchestrator destroys this VM or pod Step 2 β Let a controller scale on queued jobs Jump to heading
A controller watches for queued jobs that match your runner labels and creates runners to serve them. On Kubernetes, Actions Runner Controller does this with runner scale sets; on VMs, cloud autoscaling groups driven by webhook events or the API serve the same purpose.
# Actions Runner Controller β a runner scale set (Helm values excerpt)
githubConfigUrl: https://github.com/acme
githubConfigSecret: arc-github-app
minRunners: 2 # a small warm pool for the working day
maxRunners: 40 # hard ceiling on cost and on load to shared services
template:
spec:
containers:
- name: runner
image: registry.example.com/runners/build:2026.10
resources: { requests: { cpu: "4", memory: 8Gi } } # Workflows target the scale set by its name
jobs:
build:
runs-on: build-linux Step 3 β Win back start-up time with pre-built images Jump to heading
Ephemeral runners are only fast if starting one is fast. Bake everything a job needs that does not change per job β toolchains, language runtimes, common CLI tools, a Git mirror of large repositories β into the runner image, rebuilt on a schedule.
FROM ghcr.io/actions/actions-runner:2.319.1
RUN sudo apt-get update && sudo apt-get install -y --no-install-recommends build-essential jq git-lfs \
&& sudo rm -rf /var/lib/apt/lists/*
# A read-only reference mirror speeds up clones of the monorepo
RUN git clone --mirror https://github.com/acme/monorepo.git /opt/mirrors/monorepo.git # In the job: clone using the baked mirror as a reference, then fetch only what is new
git clone --reference-if-able /opt/mirrors/monorepo.git "$REPO_URL" work
git -C work checkout --detach "$COMMIT_SHA" A reference mirror baked into the image gives every job most of the repositoryβs objects without network transfer; the technique is covered in reusing a Git mirror on self-hosted runners.
Step 4 β Share caches without sharing state Jump to heading
Ephemeral runners cannot keep a local cache, but they can read and write a remote one: the CI providerβs cache service, an S3-compatible bucket, or a BuildKit registry cache. Keep write access scoped so jobs from pull requests cannot overwrite caches used by the default branch.
- uses: actions/cache@v4
with:
path: ~/.cache/pip
key: pip-${{ runner.os }}-${{ hashFiles('**/requirements*.txt') }} Container build caches follow the same pattern, described in caching Docker layers in CI.
Step 5 β Set limits and watch the cost Jump to heading
Autoscaling without limits turns a runaway workflow β a matrix accidentally expanded to hundreds of jobs β into a large bill and a flood of load on shared services such as package registries. Set a maximum per scale set, a minimum warm pool sized for normal mornings, and scale-down delays that avoid thrashing.
# Queue time and runner count over the day, from the CI API
gh api "orgs/$ORG/actions/runners" --jq '.runners | group_by(.status) | map({(.[0].status): length}) | add'
gh run list --limit 200 --json createdAt,startedAt \
--jq '[.[] | ((.startedAt|fromdateiso8601) - (.createdAt|fromdateiso8601))] | (add/length) | "avg queue wait: \(. | floor) s"' Validation checklist Jump to heading
Frequently Asked Questions Jump to heading
How small can the warm pool be? Jump to heading
Size it for the burst at the start of the working day, while new runners boot. Two to five runners is common for medium teams; measure queue wait at 09:00 and adjust.
Do ephemeral runners make builds slower? Jump to heading
The first minute of each job can be slower than on a warm long-lived machine. Pre-built images and remote caches recover most of that, and the reliability gain from clean state usually saves more time than it costs.
Can one scale set serve every repository? Jump to heading
It can, but separating scale sets by trust level β internal repositories, public repositories accepting outside contributions, release builds β keeps a compromise in one from reaching the others.
Related Jump to heading
- CI Caching & Runner Performance β the parent topic.
- Splitting a Long Pipeline into Parallel Jobs β more jobs make autoscaling more valuable.
- Running CI for Fork Pull Requests Safely β why untrusted code needs separate runners.
- Measuring Pipeline Duration Trends β tracking queue time alongside run time.