Measuring pipeline duration trends Jump to heading

CI gets slower the way a codebase gets messier: a few seconds at a time, each change reasonable on its own. Nobody notices a test suite growing from six minutes to seven. Twelve months later the pipeline takes eleven minutes, the merge queue backs up, and people start batching changes to avoid waiting β€” which makes conflicts and failures worse. The only defence is to measure. CI systems expose run and job timings through their APIs, so a weekly job can record them, separate the time spent waiting for a runner from the time spent running, and report medians and slow-run percentiles per workflow. With that, a regression shows up as a step in a chart the week it happens. This page builds the measurement, within CI caching and runner performance.

When to use this approach Jump to heading

  • People say CI is slow, but nobody can say how slow or since when.
  • You are about to invest in caching or runners and want a baseline to compare against.
  • A merge queue’s throughput depends on pipeline duration, as in batching pull requests to cut CI cost.
  • You want to catch slowdowns introduced by individual changes.

Step 1 β€” Separate queue time from run time Jump to heading

A pipeline’s wall-clock duration has two parts that have different causes. Queue time is waiting for a runner β€” a capacity problem. Run time is the jobs themselves β€” a pipeline problem. Measuring them together hides which one is growing.

The parts of one pipeline's durationA pipeline is created when the event arrives, waits in the queue until a runner picks up its first job, runs its jobs, and finishes. Queue time is a capacity question and run time is a pipeline question, so they are tracked separately.Createdpush event10:00:00First job startsqueue: 2m40s10:02:40Tests finishlongest job10:09:10Completedrun: 7m15s10:09:55a slow morning with a fast pipeline needs runners, not caching
# One run: created, started and completed timestamps
gh api "repos/$OWNER/$REPO/actions/runs/$RUN_ID" \
  --jq '{created: .created_at, started: .run_started_at, updated: .updated_at, conclusion}'

Step 2 β€” Collect timings for recent runs Jump to heading

Pull a few hundred recent runs of each workflow on the default branch, keep only successful ones (failed runs stop early and distort durations), and compute queue and run time for each.

#!/bin/sh
# ci-durations.sh <workflow-file> β€” TSV: date, queue seconds, run seconds
gh api --paginate "repos/$OWNER/$REPO/actions/workflows/$1/runs?branch=main&status=success&per_page=100" \
  --jq '.workflow_runs[] |
    [ .created_at[:10],
      ((.run_started_at|fromdateiso8601) - (.created_at|fromdateiso8601)),
      ((.updated_at|fromdateiso8601) - (.run_started_at|fromdateiso8601)) ] | @tsv' \
  > "durations-$1.tsv"
wc -l "durations-$1.tsv"

Restricting to the default branch gives a stable population: the same pipeline, the same triggers. Pull-request runs vary with what each change touches and are better measured separately.

Step 3 β€” Report medians and p90 per week Jump to heading

Averages are distorted by the occasional very slow run. Report the median for the typical experience and the 90th percentile for the slow end, per week.

# Weekly median and p90 of run time, in minutes
awk -F'\t' '{ cmd="date -d " $1 " +%G-W%V"; cmd | getline wk; close(cmd); print wk "\t" $3 }' durations-ci.yml.tsv |
sort | awk -F'\t' '
  { v[$1][++n[$1]] = $2 }
  END { for (w in n) { asort(v[w]); m=v[w][int(n[w]*0.5)+1]; p=v[w][int(n[w]*0.9)+1];
        printf "%s  median %.1f min  p90 %.1f min  (%d runs)\n", w, m/60, p/60, n[w] } }' | sort

This uses GNU awk’s arrays of arrays and asort; the same calculation is a few lines in any scripting language if you prefer.

Weekly median run time of the main pipelineAn illustrative quarter of weekly medians. Run time drifted slowly upward for weeks, then jumped when a new integration suite was added, and dropped after the suite was split into parallel jobs.median run time on main, minutes (illustrative)W276.8 minW307.3 minW337.9 minW3410.6 minW366.1 minthe jump in W34 traced to one pull request β€” found within a week, not a year

Step 4 β€” Find which jobs got slower Jump to heading

When the median rises, break it down by job to see where. The jobs API gives per-job start and completion times.

gh api "repos/$OWNER/$REPO/actions/runs/$RUN_ID/jobs" \
  --jq '.jobs[] | "\(.name)\t\(((.completed_at|fromdateiso8601) - (.started_at|fromdateiso8601)) / 60 | . * 10 | floor / 10) min"' |
  sort -t "$(printf '\t')" -k2 -rn

Compare the same job across a run before the regression and one after. A jump in one job, at one date, usually points to one merged change β€” find it with git log --since around that date on the files that job builds or tests.

Step 5 β€” Alert on regressions, not on noise Jump to heading

A weekly report is enough for trends. For sharper detection, alert when a week’s median exceeds the previous four weeks’ median by a margin, so random variation does not page anyone.

# Alert if this week's median is >15% above the trailing four-week median
awk -v th=1.15 '{ med[NR]=$3 } END { prev=(med[NR-1]+med[NR-2]+med[NR-3]+med[NR-4])/4;
  if (med[NR] > prev*th) printf "CI slowed: %.1f vs %.1f min\n", med[NR]/60, prev/60 }' weekly.tsv
A weekly duration reportA scheduled job pulls a week of successful runs per workflow, splits queue and run time, computes medians and the 90th percentile, compares them with the trailing four weeks, and posts a short report that flags any workflow that slowed by more than the threshold.ScheduleweeklyCollectruns API, main onlySplitqueue vs runMedian + p90per workflowReportflag > 15% jumpsone short message a week is enough to keep CI time from drifting unnoticed

Running the report on a schedule is covered in scheduled workflows and cron triggers.

Validation checklist Jump to heading

Frequently Asked Questions Jump to heading

Why exclude failed runs? Jump to heading

Failed runs stop at the first failure, so they look faster than they are, and their durations depend on where they failed. Track the failure rate separately; it matters as much as duration.

What about re-run attempts? Jump to heading

Each attempt has its own timings. Use the latest attempt for duration, and count re-runs as a separate signal β€” frequent re-runs usually mean flaky tests, as discussed in handling flaky tests in a merge queue.

Is there a target duration to aim for? Jump to heading

Under ten minutes keeps pull requests flowing for most teams; under five lets people wait for CI without switching tasks. More important than any target is that the number stops creeping up.