Writing custom secret detection rules Jump to heading

Off-the-shelf secret scanners ship hundreds of rules for public token formats β€” cloud provider keys, payment API keys, chat webhooks. They know nothing about the tokens your own systems issue: the internal service-to-service key, the signed session secret, the database DSN format your platform team invented. Those are exactly the credentials most likely to be pasted into a config file, because developers handle them daily. A custom rule closes that gap, but a careless one floods the team with false positives until everyone ignores the scanner. This page shows how to write rules that are specific enough to be trusted, and how to test them before they reach anyone’s pre-commit hook. It sits within secret scanning and remediation.

When to use this approach Jump to heading

  • Your organisation issues its own API keys, tokens or connection strings.
  • A leaked internal credential was found by a person rather than the scanner.
  • The scanner is already running β€” chosen as described in choosing a secret scanner for Git history β€” and you want to extend it rather than replace it.
  • You can collect a handful of real (revoked) examples of each credential format for testing.

Step 1 β€” Make your tokens recognisable, if you issue them Jump to heading

The best custom rule is one you barely need to write, because the token announces itself. If your platform team can change the token format, give it a fixed prefix, a known length and a checksum. Then the rule is nearly free of false positives.

Token formats that are easy and hard to detectA random 32-character hex string looks like any hash, so a rule for it either misses leaks or flags every checksum in the repository. A token with a distinctive prefix, fixed length and checksum can be matched precisely and verified offline.Unstructured tokenPrefixed + checksummedexample9f86d081884c7d65…acme_live_7Hq…_c4f1regex precisionmatches any hashmatches only tokensfalse positivesmanynear zerooffline validationimpossiblechecksum verifiesdesigning the token is the cheapest detection work you will ever do
acme_live_<30 base62 characters>_<4 hex checksum>
acme_test_<30 base62 characters>_<4 hex checksum>

Step 2 β€” Write the rule with a tight pattern and keywords Jump to heading

Most scanners use a TOML or YAML rule format with a regex, optional keywords that must appear nearby, and an entropy threshold. The example uses gitleaks syntax; other scanners have direct equivalents.

# .gitleaks.toml
[extend]
useDefault = true

[[rules]]
id = "acme-api-key"
description = "Acme internal API key"
regex = '''\bacme_(live|test)_[0-9A-Za-z]{30}_[0-9a-f]{4}\b'''
keywords = ["acme_live_", "acme_test_"]
tags = ["internal", "api-key"]

[[rules]]
id = "acme-db-dsn"
description = "Acme database DSN with embedded password"
regex = '''acmedb://[a-z0-9_-]+:([^@\s]{12,})@[a-z0-9.-]+'''
secretGroup = 1
entropy = 3.5
keywords = ["acmedb://"]

Keywords are a performance feature: the scanner skips the regex entirely unless a keyword appears, which matters when scanning full history. secretGroup tells the scanner which part of the match is the secret, so findings and redactions show the password rather than the whole DSN. The entropy threshold rejects placeholders like changeme123456.

Step 3 β€” Test rules against samples before rolling them out Jump to heading

Keep a small corpus of true positives (revoked real tokens, or generated ones in the right format) and true negatives (hashes, UUIDs, placeholders, documentation examples). Run the scanner against both and require exact results.

mkdir -p rules-test/positive rules-test/negative
printf 'API_KEY=acme_live_%s_%s\n' "$(head -c 30 /dev/urandom | base64 | tr -dc 'A-Za-z0-9' | head -c 30)" "a1b2" \
  > rules-test/positive/env.txt
printf 'sha=9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08\nkey=acme_live_EXAMPLE\n' \
  > rules-test/negative/hashes.txt

gitleaks detect --no-git --source rules-test/positive --config .gitleaks.toml --report-format json --report-path pos.json || true
gitleaks detect --no-git --source rules-test/negative --config .gitleaks.toml --report-format json --report-path neg.json || true
jq length pos.json neg.json    # expect: positives > 0, negatives == 0
A rule's path from draft to enforcementA draft rule is tested against a corpus of positives and negatives, then run across full history in report-only mode to measure noise. Only when the history scan shows real findings and almost no false positives does it move into pre-commit hooks and CI blocking.Draft ruleregex + keywordsCorpus testpos > 0, neg = 0History scanreport onlyEnforcepre-commit + CIthe history scan is where you discover the false positives the corpus missed

Step 4 β€” Measure noise across full history Jump to heading

A rule that is clean on a synthetic corpus can still match things in real code: test fixtures, generated clients, vendored libraries. Scan the full history of a few large repositories in report-only mode and look at every finding.

gitleaks detect --source . --config .gitleaks.toml --report-format json --report-path history.json --log-opts="--all"
jq -r '.[] | select(.RuleID|startswith("acme-")) | "\(.RuleID)\t\(.File)\t\(.Commit[:8])"' history.json | sort | uniq -c | sort -rn | head

Findings in test fixtures are best handled with path allow-lists that stay narrow, as described in allowlisting test fixtures without blinding the scanner. Findings anywhere else are either real leaks β€” which go into the backlog β€” or a sign the pattern is too broad.

Noise from three drafts of the same ruleAn illustrative history scan of one repository. A bare 32-character pattern matched hundreds of hashes. Adding the prefix cut that to a few dozen fixtures. Adding the checksum suffix and keywords left only the genuine leaks.findings across full history (illustrative)[0-9a-f]{32}412prefix only37prefix + checksum3three real leaks were in there all along β€” the first draft would have buried them

Step 5 β€” Ship the rules everywhere the scanner runs Jump to heading

Custom rules are only useful if every scanning point loads them: developer pre-commit hooks, CI on pull requests, scheduled history scans and the forge’s push protection if it supports custom patterns.

# .pre-commit-config.yaml β€” the hook reads .gitleaks.toml from the repo root
repos:
  - repo: https://github.com/gitleaks/gitleaks
    rev: v8.18.4
    hooks: [ { id: gitleaks } ]

Store the shared rules in one place and pull them into each repository, so updating a rule updates every scanner. The pre-push side is in blocking secrets with a pre-push scan.

Validation checklist Jump to heading

Frequently Asked Questions Jump to heading

How do I detect secrets that have no recognisable format? Jump to heading

Look for context rather than shape: assignments to variables named like password, secret or token with a high-entropy value. These generic rules are noisier, so pair them with tight entropy thresholds and keep them out of blocking hooks until they have proved themselves.

Should custom rules live in each repository? Jump to heading

Keep the source in one shared repository and distribute it, by copying in CI or by referencing a shared config file. Rules that drift between repositories mean a leak blocked in one place is allowed in another.

Can I validate whether a detected key is live? Jump to heading

If your token has a checksum, the scanner can confirm it is well-formed offline. Checking whether it is still active needs a call to the issuing service; many teams expose an internal endpoint that reports status without revealing the token’s permissions.