Forge API & Webhook Automation Jump to heading
A large part of a team’s Git workflow does not live in the repository at all. Pull requests, reviews, labels, branch protection, required checks, deploy keys, team permissions, issues and releases all live on the forge — GitHub, GitLab, Gitea, Bitbucket — and are reachable only through its interface or its API. Automation that stops at the repository boundary leaves those pieces to clicking, which does not scale past a handful of repositories and leaves no record. The forge’s API lets scripts read and change that state; its webhooks let services react to events as they happen. Both have properties automation must respect: authentication with the narrowest possible scope, rate limits that throttle careless loops, pagination that silently truncates results, and webhook deliveries that must be verified before they are trusted. This part of Git Automation & CI/CD Hook Engineering covers those foundations and the patterns built on them.
Prerequisites Jump to heading
The Two Directions of Forge Automation Jump to heading
Automation talks to a forge in two directions. Outbound, scripts call the API to read state or make changes: list repositories, update settings, open pull requests, post comments. Inbound, the forge calls your service through webhooks when something happens: a push, a pull request opened, a review submitted. Most real automation combines both — a webhook announces an event, and the handler calls the API to act on it.
Step 1 — Authenticate with the Narrowest Identity Jump to heading
Personal access tokens are convenient and almost always too broad: they act as a person, with every permission that person has, across every organisation they belong to. For automation, prefer an identity built for it.
# Short-lived installation token for a GitHub App, minted in CI
gh api -X POST "app/installations/$INSTALLATION_ID/access_tokens" \
-H "Authorization: Bearer $APP_JWT" --jq .token Whatever identity you use, give it only the permissions the job needs, scoped to the repositories it touches. The same reasoning applies to workflow tokens, covered in limiting workflow permissions per job.
Step 2 — Read Lists Completely: Pagination Jump to heading
API endpoints that return lists return pages — commonly thirty or a hundred items. A script that reads only the first page silently processes a fraction of the data, and nothing errors. Always paginate.
# gh follows pagination for you
gh api --paginate "orgs/$ORG/repos?per_page=100" --jq '.[].full_name' | wc -l
# With curl, follow the Link header until there is no rel="next"
url="https://api.github.com/orgs/$ORG/repos?per_page=100"
while [ -n "$url" ]; do
curl -fsS -D headers.txt -H "Authorization: Bearer $TOKEN" "$url" | jq -r '.[].full_name'
url=$(sed -n 's/.*<\([^>]*\)>; rel="next".*/\1/p' headers.txt)
done GraphQL APIs paginate with cursors instead; the principle is the same: loop until the response says there is no next page.
Step 3 — Respect Rate Limits Jump to heading
Every forge limits how many API requests an identity can make. Scripts that loop over hundreds of repositories with several calls each reach the limit quickly, and then either fail midway or — worse — get temporarily blocked. Read the limit headers, slow down before hitting zero, and back off when told to.
gh api rate_limit --jq '.resources.core | "\(.remaining)/\(.limit), resets \(.reset | todate)"' Handling limits properly — including secondary limits that apply to bursts of writes — is the subject of handling API rate limits in Git automation.
Step 4 — Verify Every Webhook Before Trusting It Jump to heading
A webhook endpoint is a public URL. Anyone who finds it can send it requests that look like forge events. Forges sign each delivery with a shared secret; the receiver must verify that signature before acting on the payload, using a constant-time comparison over the raw body.
import hmac, hashlib
def verified(secret: bytes, body: bytes, header: str) -> bool:
expected = "sha256=" + hmac.new(secret, body, hashlib.sha256).hexdigest()
return hmac.compare_digest(expected, header or "") The full receiver, including replay protection and per-forge differences, is in verifying webhook signatures.
Step 5 — Change Many Repositories Safely Jump to heading
The API’s biggest payoff is consistency across many repositories: the same branch protection, merge settings, labels and webhooks everywhere. The risk scales with the payoff — a mistaken loop changes everything at once. Always dry-run, change in small batches, and record what changed.
gh repo list "$ORG" --limit 1000 --json nameWithOwner,isArchived \
--jq '.[] | select(.isArchived|not) | .nameWithOwner' > repos.txt
while read -r repo; do
current=$(gh api "repos/$repo" --jq .delete_branch_on_merge)
[ "$current" = true ] || echo "would enable delete_branch_on_merge on $repo"
done < repos.txt The full pattern — dry runs, batches, change logs and rollback — is in bulk updating repository settings with the API.
Step 6 — Subscribe to the Events You Handle, Nothing More Jump to heading
Webhooks can be configured per repository, per organisation or as an app subscription, and each can receive dozens of event types. Subscribing to everything is tempting during development and expensive in production: every delivery costs the receiver a verification, a parse and usually a log line, and broad subscriptions make it harder to see which events actually drive behaviour.
Choose the narrowest scope and the smallest event list that covers the automation. A labelling bot needs pull_request with the opened and synchronize actions; a deploy trigger needs push to particular branches, or better, workflow_run completions; an audit trail may need member and repository events at organisation level.
# Inspect a repository webhook's event list and recent deliveries
gh api "repos/$OWNER/$REPO/hooks" --jq '.[] | {id, events, active, url: .config.url}'
gh api "repos/$OWNER/$REPO/hooks/$HOOK_ID/deliveries?per_page=10" \
--jq '.[] | "\(.delivered_at) \(.event)/\(.action // "-") \(.status_code)"' Within the handler, filter again by event, action and repository before doing any work. The delivery log above is also the first place to look when automation seems not to run: it shows whether the forge sent the event and what the receiver answered.
Step 7 — Design Handlers to Be Idempotent and Fast Jump to heading
Forges retry deliveries that time out or fail, and they do not guarantee ordering. A handler that takes thirty seconds will be retried while still running; a handler that assumes opened arrives before synchronize will occasionally be wrong. Two rules prevent most problems: acknowledge quickly and do the work asynchronously, and make every action safe to repeat — set a label rather than toggle it, upsert a comment rather than append one, check the current state through the API before changing it.
# Idempotent: compute the desired state, compare, change only if different
labels = {l["name"] for l in api_get(f"repos/{repo}/issues/{pr}/labels")}
wanted = labels | {"needs-review"}
if wanted != labels:
api_put(f"repos/{repo}/issues/{pr}/labels", {"labels": sorted(wanted)}) The same principle — key actions on stable identifiers, compare before writing — is what makes the bulk changes in Step 5 and the drift reports built on them safe to run repeatedly.
Integration with Adjacent Workflows Jump to heading
Forge automation connects to most other topics on this site; the boundaries are worth stating.
- Pull request bots — labelling, reviewer assignment and slash commands in pull request automation and bots are applications of the API and webhooks described here.
- Access control — permissions as code, in managing repository permissions as code, is bulk API automation with a reconcile loop.
- Server-side hooks — on self-hosted servers, hooks react to pushes; on hosted forges, webhooks fill that role, as contrasted in post-receive hooks for notifications and deploys.
- Repository scripting — plumbing answers questions about Git data; the API answers questions about forge data. Many tools need both.
Team Rollout Jump to heading
Configuration Reference Jump to heading
| Concern | Mechanism | Default to |
|---|---|---|
| Identity | App / bot account, fine-grained token, PAT | App with per-installation permissions |
| Token lifetime | Installation tokens, expiring tokens | One hour, minted per job |
| Pagination | --paginate, Link header, GraphQL cursors | Always paginate lists |
| Rate limits | X-RateLimit-*, Retry-After headers | Check before loops, back off on 403/429 |
| Webhook trust | HMAC signature header + secret | Verify every delivery, constant-time |
| Webhook replay | Delivery ID header | Store recent IDs, reject repeats |
| Bulk changes | Scripted API calls | Dry run, batches, change log |
Troubleshooting Jump to heading
| Symptom | Likely cause | Fix |
|---|---|---|
| Script processes only 30 or 100 items | Missing pagination | Use --paginate or follow Link headers |
| 403 with “rate limit exceeded” | Primary limit reached | Wait for reset; spread calls; use conditional requests |
| 403/429 on bursts of writes | Secondary limit | Serialise writes; honour Retry-After |
| Webhook handler acts on forged requests | No signature verification | Verify HMAC over the raw body |
| Automation broke when someone left | Ran on their personal token | Move to an app or bot identity |
| Settings drift back after a script ran | Another tool or person manages them | Decide one owner; run a drift report |
Frequently Asked Questions Jump to heading
REST or GraphQL? Jump to heading
REST is simpler and supported by every CLI; GraphQL fetches nested data in one request and is gentler on rate limits for complex reads. Many scripts use REST for writes and GraphQL for large reads.
Can webhooks replace polling entirely? Jump to heading
For reacting to events, yes, and they are far cheaper. Keep a periodic reconcile job anyway: deliveries can fail, and a nightly check catches anything missed.
How do I test webhook handlers locally? Jump to heading
Forges can redeliver recent events from their delivery log, and tools that tunnel to localhost make the handler reachable during development. Record a few real payloads as fixtures and test the handler against them, signatures included.
Should automation use the forge CLI or raw HTTP? Jump to heading
The CLI handles authentication, pagination and output formatting, which removes whole classes of bugs. Raw HTTP is useful when you need precise control over headers, such as conditional requests for rate limits.
Related Jump to heading
- Verifying Webhook Signatures — authenticate every delivery before acting on it.
- Bulk Updating Repository Settings with the API — change many repositories with dry runs and rollback.
- Handling API Rate Limits in Git Automation — stay within primary and secondary limits.
- Pull Request Automation & Bots — bots built on these foundations.
- Scripting Git with Plumbing Commands — the repository-side counterpart.