Detecting affected projects from git diff Jump to heading

The promise of a monorepo is one place for everything; the cost is that a naive pipeline builds and tests everything on every change. A one-line fix to a billing service should not run the web app’s end-to-end suite. The fix is to work out which projects a change affects: the projects whose files changed, plus every project that depends on them. Git provides the first half directly — the changed paths between the merge base and the branch tip. The second half needs a dependency graph, which build tools such as Nx, Bazel, Turborepo and Pants maintain, or which a small file can describe. This page builds affected-project detection from those two ingredients, and shows how to keep it correct, within monorepo branch topology.

When to use this approach Jump to heading

  • Your monorepo’s CI builds and tests every project on every change.
  • Pipeline time grows with the number of projects, not with the size of changes.
  • You use path filters, but they miss changes to shared libraries — the case covered by optimizing CI triggers for path-specific changes.
  • Projects depend on each other through internal libraries.

Step 1 — Compute the changed paths against the right base Jump to heading

The base is the merge base with the target branch, not the target’s tip. Comparing against the tip includes changes others merged since the branch was created, which are not this change’s concern.

git fetch origin main
base=$(git merge-base origin/main HEAD)
git diff --name-only -z "$base" HEAD | tr '\0' '\n' > changed.txt
wc -l changed.txt

In a merge queue or after merging, compare against the previous tip of the target instead, as described in verifying a range of commits in CI, not just the tip.

Step 2 — Map paths to projects Jump to heading

Each changed file belongs to at most one project, decided by its path. Keep the mapping in one file, so build tools and scripts agree.

# projects.yml — project roots and their internal dependencies
projects:
  billing-api:   { root: services/billing,  deps: [money, auth-client] }
  export-api:    { root: services/export,   deps: [money, storage] }
  web:           { root: apps/web,           deps: [ui-kit, auth-client] }
  money:         { root: libs/money,         deps: [] }
  auth-client:   { root: libs/auth-client,   deps: [] }
  ui-kit:        { root: libs/ui-kit,        deps: [] }
  storage:       { root: libs/storage,       deps: [] }
global: [package-lock.json, .github/workflows/, tooling/]
# Directly changed projects
yq -r '.projects | to_entries[] | "\(.value.root)\t\(.key)"' projects.yml |
while IFS="$(printf '\t')" read -r root name; do
  grep -q "^$root/" changed.txt && echo "$name"
done | sort -u > direct.txt
From a diff to the projects to testChanged paths come from the diff against the merge base. Each path maps to the project that owns it. The dependency graph is walked in reverse to add every project that depends on a changed one. Changes to global files mark everything affected.git diffmerge base..HEADPath → projectprojects.yml rootsReverse depswho depends on itGlobal files?→ everythingAffected setbuild + test theseforgetting the reverse-dependency step is how library changes ship untested

Step 3 — Add every project that depends on a changed one Jump to heading

A change to libs/money affects billing-api and export-api, even though none of their files changed. Walk the dependency graph in reverse, transitively.

# affected.py — direct changes plus transitive dependents
import sys, yaml
cfg = yaml.safe_load(open("projects.yml"))
deps = {name: set(p["deps"]) for name, p in cfg["projects"].items()}
direct = set(open("direct.txt").read().split())
affected, frontier = set(direct), set(direct)
while frontier:
    frontier = {p for p, d in deps.items() if d & frontier} - affected
    affected |= frontier
print("\n".join(sorted(affected)))
python3 affected.py > affected.txt
cat affected.txt          # money → billing-api, export-api
Pipeline time with and without affected detectionAn illustrative monorepo with twelve projects. Building and testing everything takes the same long time on every change. With affected detection, a change to one app runs only that app, a library change runs the library and its dependents, and only global changes run everything.pipeline minutes by change type (illustrative)test everything38 minone app changed6 minshared library changed14 minlockfile changed38 minmost changes are the cheap kind — that is where the saving comes from

Step 4 — Treat global files as “everything” Jump to heading

Some files affect every project: the root lockfile, shared CI configuration, build tooling, compiler settings. A change to any of them must run everything, or a broken toolchain update passes CI on an empty set.

if yq -r '.global[]' projects.yml | while read -r g; do grep -q "^$g" changed.txt && echo hit; done | grep -q hit; then
  yq -r '.projects | keys[]' projects.yml > affected.txt
  echo "global change: all projects affected"
fi
How big is the affected set?If only one project's files changed and nothing depends on it, test just that project. If a shared library changed, test it and every project that depends on it. If a global file such as the lockfile or CI config changed, test everything.What did the change touch?one app's filesThat projectfast pipelinea shared libraryLibrary + dependentsreverse walklockfile / CI / toolingEverythingsafe fallbackwhen in doubt the set should grow, never shrink

Step 5 — Drive CI from the affected set Jump to heading

Turn the list into a job matrix, so each affected project builds and tests in parallel and unaffected ones do not run at all. Keep one always-running gate job for required checks, as described in required checks for path-filtered workflows.

jobs:
  affected:
    runs-on: ubuntu-latest
    outputs: { projects: "${{ steps.a.outputs.projects }}" }
    steps:
      - uses: actions/checkout@v4
        with: { fetch-depth: 0 }
      - id: a
        run: |
          sh ci/affected.sh > affected.txt
          echo "projects=$(jq -R -s -c 'split("\n") | map(select(length>0))' affected.txt)" >> "$GITHUB_OUTPUT"
  test:
    needs: affected
    if: needs.affected.outputs.projects != '[]'
    strategy: { matrix: { project: "${{ fromJSON(needs.affected.outputs.projects) }}" } }
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - run: make -C "$(yq -r ".projects.\"${{ matrix.project }}\".root" projects.yml)" test

Step 6 — Keep the graph honest Jump to heading

The dependency file only works if it is accurate. Check it against the real imports or package manifests in CI, so a new dependency added in code but not in projects.yml fails the build rather than silently skipping tests.

# Example for JavaScript workspaces: declared deps must match package.json internal deps
for p in $(yq -r '.projects | keys[]' projects.yml); do
  root=$(yq -r ".projects.\"$p\".root" projects.yml)
  [ -f "$root/package.json" ] || continue
  actual=$(jq -r '.dependencies // {} | keys[] | select(startswith("@acme/")) | sub("@acme/";"")' "$root/package.json" | sort)
  declared=$(yq -r ".projects.\"$p\".deps[]" projects.yml | sort)
  [ "$actual" = "$declared" ] || echo "dependency drift in $p"
done

Build tools such as Nx and Bazel derive the graph from code directly, which removes this step; the pattern here suits repositories without one.

Validation checklist Jump to heading

Frequently Asked Questions Jump to heading

Should we use a build tool instead of a script? Jump to heading

If your monorepo is large or polyglot, a build tool that understands the graph — and caches results — pays for itself. The script approach is a good start and makes the concepts explicit.

What about the merge queue? Jump to heading

In a merge queue, run affected detection against the queue’s base, or simply run everything; the queue is the last gate before main, and running everything there is a common, safe choice.

Can a change be missed? Jump to heading

Yes, if the graph is wrong or a dependency is indirect — a configuration file read at runtime, a generated client. Run a full build nightly on main to catch anything affected detection missed.