Measuring repository growth over time Jump to heading
Repositories rarely become slow suddenly. They grow: a few more megabytes of history each week, a directory of fixtures that doubles every quarter, tens of thousands of stale branches, one team committing build output. By the time someone complains that clones take ten minutes, the cause has been accumulating for a year and is expensive to undo. Measuring a handful of numbers regularly β pack size, tree size, file count, commit count, ref count, largest blobs β shows trends early, when a .gitignore entry or an LFS rule still fixes them cheaply. This page sets up those measurements with git-sizer and plain Git commands, stores them as a time series, and alerts on worrying trends, within large repository performance.
When to use this approach Jump to heading
- Clone, fetch or status times have been creeping up.
- You maintain a monorepo or a repository many teams commit to.
- You want to catch large binaries and generated files before they enter history permanently.
- You are planning capacity for CI runners or Git servers.
Step 1 β Choose what to measure Jump to heading
Pick numbers that map to user-visible costs. Each one slows down a different operation.
Step 2 β Collect the numbers from a fresh mirror Jump to heading
Measure a bare mirror, updated before each run, so local branches and uncollected garbage do not distort the results.
git clone --mirror https://git.example.com/org/monorepo.git /srv/metrics/monorepo.git 2>/dev/null || \
git -C /srv/metrics/monorepo.git remote update --prune
cd /srv/metrics/monorepo.git
git gc --quiet
pack_kb=$(git count-objects -v | awk '/size-pack/ {print $2}')
files=$(git ls-tree -r --name-only HEAD | wc -l)
tree_kb=$(git ls-tree -r -l HEAD | awk '{s+=$4} END {print int(s/1024)}')
commits=$(git rev-list --count HEAD)
refs=$(git for-each-ref | wc -l)
echo "$(date +%F) $pack_kb $files $tree_kb $commits $refs" Step 3 β Run git-sizer for a broader check Jump to heading
git-sizer reports many size dimensions and flags those that are unusually large, with a concern level for each.
git-sizer --verbose --json > "sizer-$(date +%F).json"
git-sizer | grep -E '\*' # lines with asterisks are flagged as concerning Step 4 β Store a time series Jump to heading
Append each run to a file in a metrics repository, or push it to your metrics system. A plain tab-separated file is enough to plot.
out=/srv/metrics/monorepo-growth.tsv
[ -f "$out" ] || printf 'date\tpack_kb\tfiles\ttree_kb\tcommits\trefs\n' > "$out"
printf '%s\t%s\t%s\t%s\t%s\t%s\n' "$(date +%F)" "$pack_kb" "$files" "$tree_kb" "$commits" "$refs" >> "$out"
tail -n 5 "$out" | column -t Step 5 β Find what caused a jump Jump to heading
When a metric jumps, find the commits that added the most data in that period.
since=2026-09-01
git rev-list --objects --since="$since" HEAD |
git cat-file --batch-check='%(objecttype) %(objectsize:disk) %(rest)' |
awk '$1=="blob" {print $2, $3}' | sort -rn | head -n 15 The full technique, including history-wide audits, is in auditing a repository for large blobs.
Step 6 β Alert on trends, not single values Jump to heading
Absolute thresholds fire too late or too often. Alert on unusual change: a weekly pack increase several times the recent average, or file count growing faster than usual.
awk -F'\t' 'NR>1 {d[NR]=$2} END {
n=NR; if (n<6) exit
avg=(d[n-1]-d[n-5])/4; last=d[n]-d[n-1]
if (avg>0 && last>3*avg) printf "pack grew %d KB this week vs %d KB/week average\n", last, avg
}' /srv/metrics/monorepo-growth.tsv Step 7 β Act while it is still cheap Jump to heading
Content that has not yet been merged can be removed by amending a branch. Content merged last week can be removed from the tip and blocked from returning, at the cost of keeping it in history. Content that has spread through a year of history needs a rewrite. Act on the alert quickly, and block recurrences with a pre-receive or CI size check.
# CI: fail a pull request that adds a blob over 5 MB
git rev-list --objects "origin/$BASE_REF..HEAD" | git cat-file --batch-check='%(objecttype) %(objectsize) %(rest)' |
awk '$1=="blob" && $2>5*1024*1024 {print "::error::" $3 " is " int($2/1048576) " MB"; bad=1} END {exit bad}' Moving existing large files is covered in migrating large binaries to Git LFS.
Validation checklist Jump to heading
Frequently Asked Questions Jump to heading
How often should we measure? Jump to heading
Weekly is enough for most repositories; daily for very active monorepos. The point is to see a jump within days of it starting.
Why measure a mirror instead of a developer clone? Jump to heading
A mirror has every ref and no local work, and garbage collection makes its pack size comparable from run to run. Developer clones vary too much.
Is steady growth a problem? Jump to heading
Not by itself. Healthy projects grow. Watch for growth that outpaces the projectβs activity, and for step changes.
Which numbers should we share with the wider team? Jump to heading
Pack size and clone time are the two people feel directly. Publish them alongside the trend line in the engineering handbook or team channel, so a sudden jump gets noticed by the people who caused it, not only by whoever maintains the metrics job.
Should the metrics job run on a CI runner? Jump to heading
A scheduled job on a long-lived machine with a persistent mirror is cheaper, because each run fetches only the weekβs changes. A CI runner would clone the whole repository every time, which is the very cost you are trying to watch.
Related Jump to heading
- Large Repository Performance β the parent topic.
- Auditing a Repository for Large Blobs β finding where size comes from.
- Speeding Up Clones with Partial Clone β coping with size already in history.
- Measuring Pipeline Duration Trends β the same trend approach for CI.