Storing and querying attestations for audits Jump to heading

Generating attestations is the visible part of provenance work. The part that pays off is answering questions with them later: which artefacts were built from the commit that introduced a vulnerability, whether a binary in production came from a reviewed commit, what the build looked like for a release from two years ago. Those questions arrive months or years after the build, often after the CI system’s logs have expired and the forge’s attestation retention has rolled over. This page sets up storage that outlives both and an index that makes the common questions one query, within build provenance and attestations.

When to use this approach Jump to heading

  • You generate attestations but have never had to retrieve one for an investigation.
  • Your audit or compliance obligations outlast your CI provider’s log and artefact retention.
  • Incident responders need to map a vulnerable commit to every artefact built from it.
  • Attestations currently live only in a registry next to the image, or only in the forge.

Step 1 — Decide where attestations live, in more than one place Jump to heading

The forge or registry copy is convenient for verification at deploy time. It is a poor archive: retention policies, repository deletion and registry garbage collection can all remove it. Keep a second, append-only copy that you control.

Three homes for an attestationThe registry copy sits next to the artefact and is what deploy-time verification reads. The forge copy is linked to the workflow run. The archive copy, in append-only object storage, is the one that still exists when an auditor asks in three years.Registrynext to the imagedeploy-time checksForgelinked to the runretention limitedArchiveappend-only bucketyears of retentionverify from the nearest copy, audit from the archive
# Append-only bucket: object lock in compliance mode, long retention
aws s3api create-bucket --bucket example-attestations --object-lock-enabled-for-bucket
aws s3api put-object-lock-configuration --bucket example-attestations \
  --object-lock-configuration '{"ObjectLockEnabled":"Enabled","Rule":{"DefaultRetention":{"Mode":"COMPLIANCE","Years":7}}}'

Any object store with immutability controls works; the property that matters is that nobody, including your own automation, can delete or overwrite an attestation once written.

Step 2 — Archive each attestation as the release job finishes Jump to heading

Copy the signed attestation bundle into the archive under a key that encodes the artefact digest. Do it in the same job that produced it, so there is no window where it exists only in the forge.

      - uses: actions/attest-build-provenance@v1
        id: attest
        with: { subject-path: "dist/app.tar.gz" }
      - run: |
          digest=$(sha256sum dist/app.tar.gz | cut -d' ' -f1)
          aws s3 cp "${{ steps.attest.outputs.bundle-path }}" \
            "s3://example-attestations/sha256/$digest/$(date +%Y%m%dT%H%M%S)-provenance.jsonl"
# Verification: the archived bundle still verifies offline
aws s3 cp "s3://example-attestations/sha256/$digest/" ./a/ --recursive
gh attestation verify dist/app.tar.gz --bundle ./a/*-provenance.jsonl --repo "$ORG/$REPO"

Step 3 — Index attestations by commit and by artefact Jump to heading

An archive keyed by digest answers “show me the attestation for this artefact”. Audits usually start from the other side: a commit. Maintain a small index with one row per attestation, written at the same time as the archive copy.

# Extract the facts worth indexing from the provenance statement
jq -r --arg d "$digest" '
  .dsseEnvelope.payload | @base64d | fromjson |
  [ $d,
    (.predicate.buildDefinition.resolvedDependencies[0].digest.gitCommit),
    (.predicate.buildDefinition.externalParameters.workflow.ref),
    (.predicate.runDetails.builder.id),
    (.predicate.runDetails.metadata.invocationId) ] | @tsv' bundle.jsonl >> index.tsv
Writing an attestation into archive and indexThe release job produces a signed bundle, copies it into the append-only archive under the artefact digest, extracts the commit, ref, builder and run ID, and appends one row to the index so both commit-first and artefact-first questions can be answered.Bundlesigned provenanceArchivesha256/<digest>/Extractcommit, ref, runIndex rowdigest ↔ committhe index is a convenience; the signed bundle in the archive is the evidence

Store the index as a table in whatever query engine your organisation already uses. A tab-separated file in the same bucket is enough for many teams and can be queried with standard tools.

Step 4 — Answer the common audit questions Jump to heading

With the archive and index in place, the questions that used to take days take a single command.

# Which artefacts were built from commits between two points in history?
git rev-list v2.3.0..v2.4.0 > commits.txt
awk -F'\t' 'NR==FNR{c[$1]=1; next} c[$2]{print $1, $2, $3}' commits.txt index.tsv

# Which commit produced the image running in production?
running=$(kubectl get pod -l app=api -o jsonpath='{.items[0].status.containerStatuses[0].imageID}' | sed 's/.*sha256://')
awk -F'\t' -v d="$running" '$1==d{print "commit:", $2, "ref:", $3}' index.tsv

# Was that commit reviewed and signed?
git log -1 --format='%G? %GS %s' "$commit"
Starting points for an audit questionAn investigation that starts from a vulnerable commit uses the index to list every artefact built from it. One that starts from a running artefact uses the digest to find its commit. One that starts from a release tag resolves the tag and does both.What does the investigation start from?a vulnerable commitCommit → artefactslist every builda running artefactDigest → committhen check reviewa release tagTag → commit → buildsboth directionstwo columns in the index — digest and commit — answer all three

Step 5 — Re-verify archived attestations periodically Jump to heading

An archive full of attestations you cannot verify is not evidence. Signing certificates expire and trust roots rotate, so verification of old bundles depends on having stored the trust material from that time. Re-verify a sample each quarter and archive the trust roots alongside.

# Quarterly: re-verify a random sample of archived bundles
aws s3 ls s3://example-attestations/sha256/ --recursive | shuf -n 20 | awk '{print $4}' |
while read -r key; do
  aws s3 cp "s3://example-attestations/$key" b.jsonl --quiet
  gh attestation verify --bundle b.jsonl --repo "$ORG/$REPO" "oci://$IMAGE@sha256:$(echo "$key" | cut -d/ -f2)" \
    >/dev/null 2>&1 && echo "ok   $key" || echo "FAIL $key"
done

Validation checklist Jump to heading

Frequently Asked Questions Jump to heading

How long should attestations be kept? Jump to heading

At least as long as the artefacts they describe can be running anywhere, plus whatever your regulatory obligations require. Seven years is common in regulated industries; for internal services, two years past the artefact’s last deployment is a reasonable floor.

Should pull-request builds be archived too? Jump to heading

No. Only artefacts that were deployed or published need provenance retained. Archiving every build adds cost and noise without answering any audit question.

What if the source repository is later deleted? Jump to heading

The attestation still names the commit hash, but you cannot inspect that commit without a copy of the repository. Archive release-tag snapshots of repositories alongside attestations, or keep a preserving mirror as described in mirroring third-party repositories for resilience.