Using Scalar for very large repositories Jump to heading

Every technique for large repositories β€” partial clone, sparse checkout, the commit-graph, multi-pack-index, fsmonitor, the untracked cache, background maintenance β€” has to be enabled separately, in the right order, on every clone. Scalar, which ships with Git, does all of it at once. scalar clone creates a blobless partial clone with cone-mode sparse checkout, sets the recommended configuration for large repositories, and registers the clone for scheduled maintenance. scalar register applies the same settings to an existing clone. This page uses Scalar for a new clone and an existing one, explains what it configured so nothing is a mystery, and covers when it is the wrong tool, within large repository performance.

When to use this approach Jump to heading

  • Developers clone a repository too large for default settings to be comfortable.
  • You want one documented command instead of a page of setup steps.
  • Clones across the team are configured inconsistently.
  • You already use pieces individually, such as speeding up clones with partial clone.

Step 1 β€” Clone with Scalar Jump to heading

scalar clone creates an enlistment: a directory containing the working tree in src/. It starts with only the top-level files checked out.

scalar version
scalar clone https://git.example.com/org/monorepo.git ~/src/monorepo
cd ~/src/monorepo/src
git sparse-checkout list             # empty: only root files are present
git sparse-checkout add services/billing libs/common

Use --full-clone to skip sparse checkout when you need the whole tree, and --no-src to put the working tree directly in the target directory.

What scalar clone sets upScalar layers several features. At the bottom, a blobless partial clone downloads commits and trees but fetches file contents on demand. Above it, cone-mode sparse checkout limits the working tree. Then indexes and caches speed up history and status. At the top, background maintenance keeps it all current.Background maintenanceprefetch, commit-graph, repackfsmonitor + untracked cachefast statuscommit-graph + multi-pack-indexfast historycone sparse checkoutsmall working treeblobless partial clonecontents on demandeach layer can be enabled by hand β€” Scalar just does all of them consistently

Step 2 β€” Register an existing clone Jump to heading

You do not need to re-clone. scalar register applies the configuration and schedules maintenance on an existing repository. It does not convert a full clone into a partial one.

cd ~/src/existing-monorepo
scalar register
scalar list                          # all registered enlistments

Step 3 β€” See what Scalar configured Jump to heading

Scalar writes ordinary Git configuration. Inspect it so you know which settings are in effect and can reproduce them elsewhere.

git config --show-origin --get-regexp '^(core|feature|fetch|gc|index|maintenance|pack|status)\.' | sort

Typical entries include core.fsmonitor, core.untrackedCache, feature.manyFiles, fetch.writeCommitGraph, index.threads and disabled automatic gc in favour of maintenance. Exact values depend on the Git version and platform.

Step 4 β€” Rely on background maintenance Jump to heading

Scalar registers the enlistment with git maintenance, which runs hourly prefetch of remote objects, commit-graph updates, loose-object packing and incremental repacks. Foreground git fetch then has little left to download.

git maintenance run --task=prefetch      # what the hourly job does
git for-each-ref refs/prefetch/ | head   # prefetched refs, not your remote-tracking branches

Prefetch updates hidden refs, so your origin/* branches change only when you fetch. Details are in scheduling git maintenance for background repacking.

Step 5 β€” Keep it healthy Jump to heading

Run diagnostics when something seems wrong, and unregister enlistments you delete, so maintenance does not try to run on missing paths.

scalar diagnose                       # collects configuration and logs into a zip for support
scalar run all                        # run all maintenance tasks now
scalar unregister ~/src/old-clone
scalar delete ~/src/abandoned         # unregisters and removes the enlistment

⚠️ SAFETY WARNING: scalar delete removes the whole enlistment directory, including uncommitted work and unpushed branches. Check git status and git log --branches --not --remotes in the enlistment first.

Is Scalar the right tool here?For a very large repository where developers work in a few areas, use scalar clone. For an existing large clone, use scalar register. For a small or medium repository, default Git settings are enough and Scalar adds little. In CI, use a shallow or blobless clone directly rather than an enlistment with background maintenance.What are you setting up?huge repo, new clonescalar clonepartial + sparsehuge repo, existing clonescalar registersettings + maintenancesmall repo or CI jobPlain gitno daemon, no scheduleScalar is for long-lived developer clones of big repositories

Step 6 β€” Make it the documented setup Jump to heading

Put the Scalar command and the usual sparse-checkout sets in the repository’s onboarding documentation, so every developer ends up with the same configuration.

# docs/getting-started.md (excerpt)
scalar clone https://git.example.com/org/monorepo.git ~/src/monorepo
cd ~/src/monorepo/src
git sparse-checkout set --cone $(cat .sparse/billing-team.txt)

Shared sparse sets per team are described in sparse checkout for large monorepos.

Step 7 β€” Measure the difference Jump to heading

Compare a default clone with a Scalar clone on the same machine: clone time, disk use and a warm git status.

time git clone https://git.example.com/org/monorepo.git /tmp/plain >/dev/null 2>&1
time scalar clone https://git.example.com/org/monorepo.git /tmp/scalar >/dev/null 2>&1
du -sh /tmp/plain /tmp/scalar
( cd /tmp/plain && git status >/dev/null && time git status >/dev/null )
( cd /tmp/scalar/src && git status >/dev/null && time git status >/dev/null )
Default clone versus Scalar cloneAn illustrative comparison on a large monorepo. The default clone downloads all history and file contents and checks out the whole tree. The Scalar clone downloads commits and trees only, checks out a small cone, and keeps status fast with a filesystem monitor.git clonescalar clonedownloadall blobs, all historycommits + treesworking treeeverythingchosen conewarm statussecondssub-secondmaintenancegc when triggeredscheduled, hourlyillustrative β€” the gap grows with repository size

Validation checklist Jump to heading

Frequently Asked Questions Jump to heading

Does Scalar need a special server? Jump to heading

No. It uses standard Git protocol features. The server must support partial clone filters, which current hosted forges do.

Can I use Git normally inside an enlistment? Jump to heading

Yes. It is an ordinary Git repository with particular settings. Every Git command works; missing file contents are fetched on demand.

What happens offline? Jump to heading

Files you have checked out are available. Operations needing contents not yet downloaded β€” checking out an old commit, blaming a file outside your cone β€” fail until you are back online.

How do I undo Scalar’s settings? Jump to heading

Run scalar unregister to stop scheduled maintenance, then remove the configuration entries it added with git config --unset. The partial clone and sparse checkout remain; git sparse-checkout disable restores the full tree, fetching missing contents as needed.

Does Scalar work on every platform? Jump to heading

It is part of Git on Linux, macOS and Windows. The filesystem monitor depends on platform support, so check git fsmonitor--daemon status after cloning; the rest works everywhere.