Triaging a backlog of historical secret findings Jump to heading

The first time a team scans the full history of its repositories, the result is rarely a handful of findings. It is hundreds or thousands: the same test key repeated in two hundred commits, a dozen expired credentials from systems decommissioned years ago, placeholder values that tripped a generic rule, and — somewhere in there — a production database password that is still valid. Working through the list from top to bottom wastes weeks on noise while the live credential stays exposed. Triage means grouping, ranking and acting on the dangerous findings first, then closing the rest with a recorded reason. This page gives that process, within secret scanning and remediation.

When to use this approach Jump to heading

  • A first scan of full history, across one repository or an organisation, returned more findings than anyone can read individually.
  • Findings arrive with no owner and no indication of whether the credential still works.
  • You need to show auditors that every historical finding was assessed, not ignored.
  • Newly introduced secrets are already blocked going forward, as in scanning pull requests for secrets in CI; this page deals with what was already there.

Step 1 — Deduplicate by secret value, not by finding Jump to heading

A single key committed once and then carried through a hundred later commits is one problem, not a hundred. Group findings by a hash of the secret itself and keep the earliest commit, the latest commit and every file it appeared in.

# From a gitleaks JSON report of full history
jq -r '.[] | [(.Secret|@base64), .RuleID, .File, .Commit, .Date] | @tsv' history.json |
awk -F'\t' '{key=$1; n[key]++; rule[key]=$2; files[key]=files[key] " " $3;
             if(!(key in first) || $5<firstd[key]){first[key]=$4; firstd[key]=$5}
             if(!(key in last)  || $5>lastd[key]) {last[key]=$4;  lastd[key]=$5}}
     END{for(k in n) print n[k] "\t" rule[k] "\t" first[k] "\t" firstd[k] "\t" last[k]}' |
sort -rn > unique-secrets.tsv
wc -l history.json unique-secrets.tsv
How a raw history scan shrinks under triageAn illustrative first scan: over a thousand raw findings collapse to around a hundred unique secrets once duplicates are grouped. Removing known placeholders and fixtures leaves a few dozen real credentials, of which a handful are still live and need urgent revocation.findings remaining after each triage stage (illustrative)raw findings1240unique secrets118real credentials41still live6the six at the bottom are the whole reason for doing this — find them first

Step 2 — Rank by exposure and by liveness Jump to heading

Two questions decide priority. Where was the secret exposed — a public repository, a widely cloned internal one, a private one with few readers? And does it still work? A live credential in a public repository is an emergency; a dead one in a private repository is a cleanup task.

Priority from exposure and livenessA credential that still works and was exposed publicly must be revoked within hours. A live credential in a private repository is urgent but can follow normal change processes. A dead credential needs only to be recorded and, optionally, removed from history.Still liveAlready deadpublic repositoryrevoke nowrecord, consider rewriteinternal, widely clonedrevoke this weekrecordprivate, few readersrevoke, schedulerecordliveness changes the urgency far more than the age of the commit

Liveness checks depend on the credential type. Many providers offer a harmless identity call; for internal tokens, ask the issuing service.

# Example liveness probes — run only from a controlled environment
aws sts get-caller-identity                                   # with the found AWS key exported
curl -s -o /dev/null -w '%{http_code}\n' -H "Authorization: Bearer $TOKEN" https://api.example.com/v1/whoami

⚠️ SAFETY WARNING: Testing a found credential is itself a use of it, and it appears in the provider’s logs as such. Run probes only from an approved security environment, record that you did it and why, and never test a credential by performing any action beyond an identity check. If you are unsure whether testing is permitted, treat the credential as live and revoke it.

Step 3 — Assign owners and revoke live credentials Jump to heading

Each live credential needs an owner who can revoke and replace it without breaking production. Use the file path and the commit author as starting points, and the service the credential belongs to as the deciding factor.

# Who touched the file where each live secret first appeared?
while IFS=$'\t' read -r count rule first date last; do
  git log -1 --format="%an <%ae>" "$first"
done < live-secrets.tsv

Revocation follows the runbook in responding to a leaked credential: revoke, replace, deploy the replacement, confirm the old one fails. Removing the value from history comes after revocation, never instead of it.

Step 4 — Close everything else with a recorded reason Jump to heading

Every finding that is not live still needs a disposition, or the next scan reports it again and nobody can tell whether it was looked at. Record each one in a register with a reason from a short, fixed list.

# secret-register.tsv
fingerprint    rule              disposition      reason                         by     date
3f9a…          aws-access-key    revoked          rotated in INC-2291            priya  2026-10-03
71c2…          generic-api-key   false-positive   example value in docs          tom    2026-10-03
a08e…          acme-db-dsn       dead             database decommissioned 2024   maria  2026-10-04
c1d7…          private-key       test-fixture     generated for unit tests       tom    2026-10-04

Feed the fingerprints of closed findings back into the scanner’s baseline or ignore file, so future scans report only new findings. Keep the register under version control; it is the evidence that the backlog was handled.

# Generate the scanner's ignore entries from the register
awk -F'\t' 'NR>1 && $3!="revoked"{print $1}' secret-register.tsv > .gitleaksignore

Step 5 — Decide whether to rewrite history Jump to heading

Revocation makes a leaked credential harmless. Rewriting history removes it from the repository, which matters mainly for public repositories, for credentials that cannot be revoked (such as a customer’s data), and for compliance regimes that require it. Rewriting is disruptive, so decide per repository rather than per secret.

Rewrite history, or not?If a credential was revoked and the repository is private, recording it is usually enough. If the repository is public, or the exposed value cannot be revoked, rewriting history is justified despite the disruption. Compliance requirements can force a rewrite regardless.Revoked — does history still need cleaning?private repo, revocableNo rewriterecord and baselinepublic repoRewritecoordinate a force-pushvalue can't be revokedRewriteplus notify ownersa rewrite never substitutes for revocation — clones made earlier still contain the secret

If you do rewrite, batch all secrets for a repository into one rewrite, following removing a leaked secret from Git history, and rotate anything the rewrite touched afterwards as described in rotating credentials after a history rewrite.

Validation checklist Jump to heading

Frequently Asked Questions Jump to heading

How long should triage of a large backlog take? Jump to heading

The live-credential pass should take days, not weeks — it is a small number of secrets once duplicates are grouped. Closing the long tail can take longer and can be shared out by repository owners, as long as the register is kept current.

Can we skip the dead credentials entirely? Jump to heading

Not without a record. An auditor, or your own future self, needs to see that a finding was examined and why it was judged dead. One line in the register is enough.

What if a credential’s owner cannot be found? Jump to heading

Revoke it anyway, from the issuing side, and watch for breakage. An unowned live credential is more dangerous than the outage its revocation might cause, and the breakage will reveal the owner quickly.