Triaging a backlog of historical secret findings Jump to heading
The first time a team scans the full history of its repositories, the result is rarely a handful of findings. It is hundreds or thousands: the same test key repeated in two hundred commits, a dozen expired credentials from systems decommissioned years ago, placeholder values that tripped a generic rule, and — somewhere in there — a production database password that is still valid. Working through the list from top to bottom wastes weeks on noise while the live credential stays exposed. Triage means grouping, ranking and acting on the dangerous findings first, then closing the rest with a recorded reason. This page gives that process, within secret scanning and remediation.
When to use this approach Jump to heading
- A first scan of full history, across one repository or an organisation, returned more findings than anyone can read individually.
- Findings arrive with no owner and no indication of whether the credential still works.
- You need to show auditors that every historical finding was assessed, not ignored.
- Newly introduced secrets are already blocked going forward, as in scanning pull requests for secrets in CI; this page deals with what was already there.
Step 1 — Deduplicate by secret value, not by finding Jump to heading
A single key committed once and then carried through a hundred later commits is one problem, not a hundred. Group findings by a hash of the secret itself and keep the earliest commit, the latest commit and every file it appeared in.
# From a gitleaks JSON report of full history
jq -r '.[] | [(.Secret|@base64), .RuleID, .File, .Commit, .Date] | @tsv' history.json |
awk -F'\t' '{key=$1; n[key]++; rule[key]=$2; files[key]=files[key] " " $3;
if(!(key in first) || $5<firstd[key]){first[key]=$4; firstd[key]=$5}
if(!(key in last) || $5>lastd[key]) {last[key]=$4; lastd[key]=$5}}
END{for(k in n) print n[k] "\t" rule[k] "\t" first[k] "\t" firstd[k] "\t" last[k]}' |
sort -rn > unique-secrets.tsv
wc -l history.json unique-secrets.tsv Step 2 — Rank by exposure and by liveness Jump to heading
Two questions decide priority. Where was the secret exposed — a public repository, a widely cloned internal one, a private one with few readers? And does it still work? A live credential in a public repository is an emergency; a dead one in a private repository is a cleanup task.
Liveness checks depend on the credential type. Many providers offer a harmless identity call; for internal tokens, ask the issuing service.
# Example liveness probes — run only from a controlled environment
aws sts get-caller-identity # with the found AWS key exported
curl -s -o /dev/null -w '%{http_code}\n' -H "Authorization: Bearer $TOKEN" https://api.example.com/v1/whoami ⚠️ SAFETY WARNING: Testing a found credential is itself a use of it, and it appears in the provider’s logs as such. Run probes only from an approved security environment, record that you did it and why, and never test a credential by performing any action beyond an identity check. If you are unsure whether testing is permitted, treat the credential as live and revoke it.
Step 3 — Assign owners and revoke live credentials Jump to heading
Each live credential needs an owner who can revoke and replace it without breaking production. Use the file path and the commit author as starting points, and the service the credential belongs to as the deciding factor.
# Who touched the file where each live secret first appeared?
while IFS=$'\t' read -r count rule first date last; do
git log -1 --format="%an <%ae>" "$first"
done < live-secrets.tsv Revocation follows the runbook in responding to a leaked credential: revoke, replace, deploy the replacement, confirm the old one fails. Removing the value from history comes after revocation, never instead of it.
Step 4 — Close everything else with a recorded reason Jump to heading
Every finding that is not live still needs a disposition, or the next scan reports it again and nobody can tell whether it was looked at. Record each one in a register with a reason from a short, fixed list.
# secret-register.tsv
fingerprint rule disposition reason by date
3f9a… aws-access-key revoked rotated in INC-2291 priya 2026-10-03
71c2… generic-api-key false-positive example value in docs tom 2026-10-03
a08e… acme-db-dsn dead database decommissioned 2024 maria 2026-10-04
c1d7… private-key test-fixture generated for unit tests tom 2026-10-04 Feed the fingerprints of closed findings back into the scanner’s baseline or ignore file, so future scans report only new findings. Keep the register under version control; it is the evidence that the backlog was handled.
# Generate the scanner's ignore entries from the register
awk -F'\t' 'NR>1 && $3!="revoked"{print $1}' secret-register.tsv > .gitleaksignore Step 5 — Decide whether to rewrite history Jump to heading
Revocation makes a leaked credential harmless. Rewriting history removes it from the repository, which matters mainly for public repositories, for credentials that cannot be revoked (such as a customer’s data), and for compliance regimes that require it. Rewriting is disruptive, so decide per repository rather than per secret.
If you do rewrite, batch all secrets for a repository into one rewrite, following removing a leaked secret from Git history, and rotate anything the rewrite touched afterwards as described in rotating credentials after a history rewrite.
Validation checklist Jump to heading
Frequently Asked Questions Jump to heading
How long should triage of a large backlog take? Jump to heading
The live-credential pass should take days, not weeks — it is a small number of secrets once duplicates are grouped. Closing the long tail can take longer and can be shared out by repository owners, as long as the register is kept current.
Can we skip the dead credentials entirely? Jump to heading
Not without a record. An auditor, or your own future self, needs to see that a finding was examined and why it was judged dead. One line in the register is enough.
What if a credential’s owner cannot be found? Jump to heading
Revoke it anyway, from the issuing side, and watch for breakage. An unowned live credential is more dangerous than the outage its revocation might cause, and the breakage will reveal the owner quickly.
Related Jump to heading
- Secret Scanning & Remediation — the parent topic.
- Writing Custom Secret Detection Rules — the rules that produce your history findings.
- Allowlisting Test Fixtures Without Blinding the Scanner — handling the fixture findings properly.
- Auditing a Repository for Large Blobs — another history-wide audit with the same triage shape.