Rejecting large files before push Jump to heading

A 300 MB video, a database dump or a build artefact committed by mistake becomes part of every clone forever, even if the next commit deletes it. Server-side limits reject the push, but only after uploading the whole pack, and some servers allow files well above what you would want in history. A pre-push hook can catch the file before anything is sent. The catch is that it must check the right thing: not the files present at the tip, but every blob introduced by every commit in the push — because the file deleted in the last commit is still in the second-to-last, and the push sends both. This page builds a check that does that, suggests Git LFS for files that should be large, and stays fast, within pre-push validation rules.

When to use this approach Jump to heading

  • Large files have been committed by accident and had to be removed with a history rewrite.
  • Your server limits file size, and you want developers to find out before uploading.
  • Some file types legitimately need to be large and should go through Git LFS instead.
  • You already compute pushed commits with the helper from finding the new commits in a pre-push hook.

Step 1 — Check blobs in every pushed commit, not just the tip Jump to heading

Listing files at HEAD misses a file added and deleted within the push. Instead, list every object reachable from the pushed commits that the remote does not have, and check the size of each blob.

#!/bin/sh
# scripts/hooks/check-large-blobs.sh <limit-bytes> — reads pushed commit IDs on stdin
set -eu
limit=${1:-5242880}                          # 5 MiB default
commits=$(cat)
[ -n "$commits" ] || exit 0
git rev-list --objects $commits --not --remotes |
  git cat-file --batch-check='%(objecttype) %(objectname) %(objectsize) %(rest)' |
  awk -v lim="$limit" '$1=="blob" && $3>lim { printf "✗ %s (%.1f MiB)\n", $4, $3/1048576; bad=1 } END { exit bad }'
Why checking the tip is not enoughCommit C1 adds a 300 MB dump, and commit C2 deletes it. A check of the tip sees no large file. The push still sends C1's blob, because history contains it. Checking every new blob catches the file in C1.the tip is clean; the push is notpushed commitsC1C2large blobdump.sqlat the tipdeletedrev-list --objects over the pushed range sees every blob the push will upload Why checking the tip is not enoughCommit C1 adds a 300 MB dump, and commit C2 deletes it. A check of the tip sees no large file. The push still sends C1's blob, because history contains it. Checking every new blob catches the file in C1.the tip is clean; the push is notpushed commitsC1C2large blobdump.sqlat the tipdeletedrev-list --objects over the pushed range sees every blob the push will upload

--not --remotes limits the walk to objects the remote does not already have, which keeps the check fast even on large repositories: only new objects are examined.

Step 2 — Wire it into the pre-push hook Jump to heading

Use the shared commit computation, then pass the list to the size check. Read the limit from Git configuration so each repository can tune it.

# .husky/pre-push
input=$(cat)
commits=$(printf '%s\n' "$input" | sh scripts/hooks/pushed-commits.sh "$1")
limit=$(git config --get hooks.maxBlobSize || echo 5242880)
printf '%s\n' "$commits" | sh scripts/hooks/check-large-blobs.sh "$limit" || {
  echo "Large files would be pushed. Remove them from history, or track them with Git LFS."
  exit 1
}
# Verification: a 10 MiB file added and removed in two commits is still caught
head -c 10485760 /dev/urandom > big.bin && git add big.bin && git commit -qm "add big"
git rm -q big.bin && git commit -qm "remove big"
git rev-list @{u}..HEAD | sh scripts/hooks/check-large-blobs.sh 5242880; echo "exit: $?"   # 1
Time to check a push for large blobsAn illustrative repository with a long history. Walking only the objects the remote lacks keeps a typical push check under a tenth of a second. Walking every object reachable from the tip would take far longer, which is why --not --remotes matters.seconds to check one push (illustrative)new objects only0.08 sfirst push of old branch1.2 sall reachable objects34 sthe hook pays for what the push sends, not for the repository size Time to check a push for large blobsAn illustrative repository with a long history. Walking only the objects the remote lacks keeps a typical push check under a tenth of a second. Walking every object reachable from the tip would take far longer, which is why --not --remotes matters.seconds to check one push (illustrative)new objects only0.08 sfirst push of old branch1.2 sall reachable objects34 sthe hook pays for what the push sends, not for the repository size

Step 3 — Remove the file from unpushed history Jump to heading

Because the commits have not been pushed, removing the file is a local rewrite with no effect on anyone else. Drop the file from the commit that added it with an interactive rebase.

git log --oneline --diff-filter=A -- big.bin           # which commit added it
git rebase -i @{upstream}                               # mark that commit 'edit'
git rm --cached big.bin && git commit --amend --no-edit
git rebase --continue

⚠️ SAFETY WARNING: This rewrites local commits that have not been pushed — safe for you, but make sure none of them exist on the remote. If the large file was already pushed in an earlier push, removing it needs a full history rewrite and coordination with everyone who fetched it, as described in rewriting history with git filter-repo.

Step 4 — Route legitimately large files through Git LFS Jump to heading

Some large files belong in the repository: design assets, test fixtures, model weights. Track them with Git LFS so history stores small pointers instead of the content. LFS pointer files are tiny, so they pass the size check automatically.

git lfs install
git lfs track '*.psd' 'fixtures/*.parquet'
git add .gitattributes
git add design/hero.psd && git commit -m "chore: track design source via LFS"
git lfs ls-files
What to do with a large fileIf the file was committed by accident, remove it from the unpushed history. If it is a generated artefact, add it to gitignore and publish it elsewhere. If it genuinely belongs with the code, track its pattern with Git LFS so history holds only a pointer.Why is the file in the commit?accidentRemove from historyrebase -i, editbuild outputIgnore + publishartefact storebelongs with codeGit LFSsmall pointer in historythe hook's message should name all three options What to do with a large fileIf the file was committed by accident, remove it from the unpushed history. If it is a generated artefact, add it to gitignore and publish it elsewhere. If it genuinely belongs with the code, track its pattern with Git LFS so history holds only a pointer.Why is the file in the commit?accidentRemove from historyrebase -i, editbuild outputIgnore + publishartefact storebelongs with codeGit LFSsmall pointer in historythe hook's message should name all three options

Migrating files already in history to LFS is covered in migrating large binaries to Git LFS.

Step 5 — Back it up on the server Jump to heading

Hooks can be skipped. The server should enforce a hard limit too, ideally the same one, so the client hook is simply the early warning. Self-hosted servers can use a pre-receive check; hosted forges have built-in limits and push rules.

# Server-side equivalent in a pre-receive hook (new objects only)
git rev-list --objects "$new" --not --all |
  git cat-file --batch-check='%(objecttype) %(objectsize) %(rest)' |
  awk -v lim=5242880 '$1=="blob" && $2>lim { print "rejected large file:", $3; bad=1 } END { exit bad }'

The server side is covered in enforcing file size limits on the remote.

Validation checklist Jump to heading

Frequently Asked Questions Jump to heading

What limit should we choose? Jump to heading

Small enough that accidents are caught and large enough that ordinary assets pass: a few megabytes suits most code repositories. Forges commonly warn above 50 MB and refuse above 100 MB, which is far too high to be useful as the only limit.

Does the check slow down every push? Jump to heading

It walks only new objects, so a normal push with a handful of commits takes milliseconds. A first push of a large existing branch takes longer, once.

What about files that are large but compress well? Jump to heading

The check uses the object’s uncompressed size, which is what clones eventually materialise. Large text files such as generated JSON still deserve scrutiny even if they compress.