Measuring repository growth over time Jump to heading

Repositories rarely become slow suddenly. They grow: a few more megabytes of history each week, a directory of fixtures that doubles every quarter, tens of thousands of stale branches, one team committing build output. By the time someone complains that clones take ten minutes, the cause has been accumulating for a year and is expensive to undo. Measuring a handful of numbers regularly β€” pack size, tree size, file count, commit count, ref count, largest blobs β€” shows trends early, when a .gitignore entry or an LFS rule still fixes them cheaply. This page sets up those measurements with git-sizer and plain Git commands, stores them as a time series, and alerts on worrying trends, within large repository performance.

When to use this approach Jump to heading

  • Clone, fetch or status times have been creeping up.
  • You maintain a monorepo or a repository many teams commit to.
  • You want to catch large binaries and generated files before they enter history permanently.
  • You are planning capacity for CI runners or Git servers.

Step 1 β€” Choose what to measure Jump to heading

Pick numbers that map to user-visible costs. Each one slows down a different operation.

Growth metrics and what they slow downPack size drives clone and fetch time. Checkout size and file count drive status and checkout time. Commit count drives history walks. Ref count drives fetch negotiation and ref advertisement. The largest blobs show where size comes from.Pack sizeclone, fetchFiles at HEADstatus, checkoutCommitslog, merge-baseRefsfetch negotiationLargest blobswhere size comes fromsix numbers, collected weekly, are enough to see trouble coming

Step 2 β€” Collect the numbers from a fresh mirror Jump to heading

Measure a bare mirror, updated before each run, so local branches and uncollected garbage do not distort the results.

git clone --mirror https://git.example.com/org/monorepo.git /srv/metrics/monorepo.git 2>/dev/null || \
  git -C /srv/metrics/monorepo.git remote update --prune
cd /srv/metrics/monorepo.git
git gc --quiet
pack_kb=$(git count-objects -v | awk '/size-pack/ {print $2}')
files=$(git ls-tree -r --name-only HEAD | wc -l)
tree_kb=$(git ls-tree -r -l HEAD | awk '{s+=$4} END {print int(s/1024)}')
commits=$(git rev-list --count HEAD)
refs=$(git for-each-ref | wc -l)
echo "$(date +%F) $pack_kb $files $tree_kb $commits $refs"

Step 3 β€” Run git-sizer for a broader check Jump to heading

git-sizer reports many size dimensions and flags those that are unusually large, with a concern level for each.

git-sizer --verbose --json > "sizer-$(date +%F).json"
git-sizer | grep -E '\*'          # lines with asterisks are flagged as concerning

Step 4 β€” Store a time series Jump to heading

Append each run to a file in a metrics repository, or push it to your metrics system. A plain tab-separated file is enough to plot.

out=/srv/metrics/monorepo-growth.tsv
[ -f "$out" ] || printf 'date\tpack_kb\tfiles\ttree_kb\tcommits\trefs\n' > "$out"
printf '%s\t%s\t%s\t%s\t%s\t%s\n' "$(date +%F)" "$pack_kb" "$files" "$tree_kb" "$commits" "$refs" >> "$out"
tail -n 5 "$out" | column -t
Pack size by quarterAn illustrative repository whose pack size grows steadily for three quarters, then jumps when a team starts committing generated test fixtures. The jump is visible in the weekly series the week it starts, long before anyone notices slower clones.pack size in MB (illustrative)Q1820 MBQ2910 MBQ31010 MBQ41940 MBa step change, not steady growth, is what deserves a same-week investigation

Step 5 β€” Find what caused a jump Jump to heading

When a metric jumps, find the commits that added the most data in that period.

since=2026-09-01
git rev-list --objects --since="$since" HEAD |
  git cat-file --batch-check='%(objecttype) %(objectsize:disk) %(rest)' |
  awk '$1=="blob" {print $2, $3}' | sort -rn | head -n 15

The full technique, including history-wide audits, is in auditing a repository for large blobs.

Absolute thresholds fire too late or too often. Alert on unusual change: a weekly pack increase several times the recent average, or file count growing faster than usual.

awk -F'\t' 'NR>1 {d[NR]=$2} END {
  n=NR; if (n<6) exit
  avg=(d[n-1]-d[n-5])/4; last=d[n]-d[n-1]
  if (avg>0 && last>3*avg) printf "pack grew %d KB this week vs %d KB/week average\n", last, avg
}' /srv/metrics/monorepo-growth.tsv
A metric jumped β€” what next?If pack size jumped, find the largest new blobs and move them to LFS or an artefact store. If files at HEAD jumped, look for generated or vendored directories and ignore them. If refs jumped, look for automation creating branches or tags that are never deleted.Which metric jumped?pack sizeLargest new blobsLFS / artefact storefiles at HEADGenerated / vendored dirsignore, removerefsAutomation branchesdelete merged refsmost jumps trace back to one commit or one bot

Step 7 β€” Act while it is still cheap Jump to heading

Content that has not yet been merged can be removed by amending a branch. Content merged last week can be removed from the tip and blocked from returning, at the cost of keeping it in history. Content that has spread through a year of history needs a rewrite. Act on the alert quickly, and block recurrences with a pre-receive or CI size check.

# CI: fail a pull request that adds a blob over 5 MB
git rev-list --objects "origin/$BASE_REF..HEAD" | git cat-file --batch-check='%(objecttype) %(objectsize) %(rest)' |
  awk '$1=="blob" && $2>5*1024*1024 {print "::error::" $3 " is " int($2/1048576) " MB"; bad=1} END {exit bad}'

Moving existing large files is covered in migrating large binaries to Git LFS.

Validation checklist Jump to heading

Frequently Asked Questions Jump to heading

How often should we measure? Jump to heading

Weekly is enough for most repositories; daily for very active monorepos. The point is to see a jump within days of it starting.

Why measure a mirror instead of a developer clone? Jump to heading

A mirror has every ref and no local work, and garbage collection makes its pack size comparable from run to run. Developer clones vary too much.

Is steady growth a problem? Jump to heading

Not by itself. Healthy projects grow. Watch for growth that outpaces the project’s activity, and for step changes.

Which numbers should we share with the wider team? Jump to heading

Pack size and clone time are the two people feel directly. Publish them alongside the trend line in the engineering handbook or team channel, so a sudden jump gets noticed by the people who caused it, not only by whoever maintains the metrics job.

Should the metrics job run on a CI runner? Jump to heading

A scheduled job on a long-lived machine with a persistent mirror is cheaper, because each run fetches only the week’s changes. A CI runner would clone the whole repository every time, which is the very cost you are trying to watch.