TL;DR: HMAC signature verification proves a webhook wasn't tampered with and came from a holder of the shared secret — it says nothing about when the request was sent or whether you've already processed it. Real webhook security needs a timestamp tolerance window, event-ID deduplication, a key rotation plan with overlap, and SSRF-hardened outbound delivery if you also send webhooks. This is a design guide, not a "add HMAC and ship it" checklist.

Most teams treat webhook security as a solved problem the moment they add signature verification. Compute an HMAC-SHA256 over the payload, compare it to the header, done. That's necessary, but it's roughly a third of the actual threat model. Signature verification alone is vulnerable to replay, gives you no story for what happens when a secret leaks, and if you're the one sending webhooks to customer-supplied URLs, it says nothing about the server-side request forgery risk sitting in your delivery path.

This piece works through the full surface: what HMAC verification actually proves, why timestamps and idempotency keys are not optional, how key rotation should work without a synchronized flag day, and what changes if you're building a webhook sender rather than just a receiver. We'll reference the Standard Webhooks specification, which OpenAI, Anthropic, Twilio, Kong, PagerDuty, and Supabase have converged on, because reinventing this from scratch is rarely worth it. If retry logic and dead-letter handling for at-least-once messaging is the gap in your stack rather than security, that's a separate design problem worth reading up on directly.

What HMAC signature verification actually proves — and doesn't

A provider signs a webhook by hashing the payload (and usually a timestamp and ID) with a shared secret using HMAC-SHA256, then sends that signature in a header. Your receiver recomputes the same hash and compares it, using a constant-time comparison function to avoid timing side-channels.

This proves two things: the payload wasn't altered in transit, and the sender possesses the shared secret. It proves nothing about:

  • Freshness. A captured, valid request is just as valid the tenth time it's replayed.
  • Uniqueness. Nothing stops the same event from being processed twice if your retry handling or network hiccups cause a duplicate delivery.
  • Trust boundaries downstream. If your handler forwards the payload to another internal service without re-validating, that service just inherited trust from a header check it never performed.

The first mistake teams make is re-serializing the JSON body before hashing it — pretty-printing it, reordering keys, or parsing and re-stringifying. Any of that changes the byte sequence and breaks the signature. Always verify against the raw request body bytes, before your framework's JSON middleware has touched them. In Express this typically means registering a raw body parser scoped to the webhook route before the global JSON parser runs; in most frameworks it means capturing the body as a buffer, not letting the router deserialize it first.

Replay protection: timestamp tolerance plus deduplication

A signature only proves authenticity at the moment of signing. Stopping replay requires two additional, complementary checks.

Timestamp tolerance. The signing scheme should include a timestamp in the signed content (Standard Webhooks calls this webhook-timestamp; Stripe's Stripe-Signature header embeds it the same way). Your receiver rejects anything outside a tolerance window — five minutes is the de facto industry default, used by both Stripe and Svix's libraries, and it's what the Standard Webhooks spec recommends while leaving the exact value to implementers. Ten minutes is defensible if your infrastructure has meaningful clock drift; anything much wider erodes the point of having a window at all.

Idempotency on event ID. The tolerance window still leaves several minutes during which a captured request replays successfully. Close that gap by treating the delivery ID as an idempotency key — Stripe's evt_ ID, GitHub's X-GitHub-Delivery GUID, Shopify's X-Shopify-Webhook-Id, or Standard Webhooks' webhook-id. Remember IDs you've already processed and reject repeats.

The critical implementation detail: make the claim atomic, not check-then-act.

sql
-- Atomic claim, not a SELECT followed by an INSERT
INSERT INTO processed_webhooks (event_id, received_at)
VALUES ($1, now())
ON CONFLICT (event_id) DO NOTHING;
-- 0 rows affected = already processed, skip
-- 1 row affected = first time, proceed

A SELECT to check existence followed by a separate INSERT has a race window — two concurrent workers can both see "not present" and both proceed. INSERT ... ON CONFLICT DO NOTHING (Postgres/SQLite) or SET NX EX (Redis) closes it in a single round-trip.

On TTL: key the retention to the sender's actual retry window plus your own queue lag and clock skew, not a round number picked at random. Stripe retries idempotency keys for at least 24 hours; a common approach is a 48-hour TTL to comfortably cover most providers' retry schedules with margin. Key on the event ID, never the resource ID — keying on the resource blocks legitimate future updates to that same resource from being processed as "new."

LayerProtects againstTypical implementation
HMAC signatureTampering, spoofed senderHMAC-SHA256 over id.timestamp.payload, constant-time compare
Timestamp toleranceLong-delay replayReject if \|now - timestamp\| > 300s
Event-ID dedupReplay within the tolerance window, duplicate deliveryAtomic insert with unique constraint, 24-72h TTL
Constant-time compareTiming side-channel on signature checkhmac.compare_digest / crypto.timingSafeEqual, never ==

Key rotation without a flag day

Signing secrets get rotated for two reasons: scheduled hygiene (commonly every 90 days) and incident response (a secret leaked and needs to die now). Both fail badly if your rotation plan is "update the secret everywhere at the same instant," because webhook senders and receivers are never perfectly synchronized — in-flight requests signed with the old secret will still be arriving after you've swapped to the new one.

The pattern that avoids downtime is dual-secret validation with an overlap window:

  1. Generate the new secret; the sender starts signing with both the old and new secret simultaneously (multiple v1= entries in one signature header, as Stripe does during dashboard rotation).
  2. The receiver accepts a signature that matches either secret during the overlap period.
  3. After the overlap window closes (Stripe's dashboard allows up to 24 hours), the old secret is retired and removed from the receiver's accepted set.
python
def verify_signature(payload: bytes, header_sig: str, secrets: list[str]) -> bool:
    for secret in secrets:  # e.g. [new_secret, old_secret] during rotation
        expected = hmac.new(secret.encode(), payload, hashlib.sha256).hexdigest()
        if hmac.compare_digest(expected, header_sig):
            return True
    return False

For incident-driven rotation — a secret leaked in a log file or a public repo — skip the overlap and revoke immediately; the cost of a brief delivery gap is lower than the cost of a known-compromised secret staying valid. Either way, store secrets in a dedicated secrets manager (Vault, AWS Secrets Manager, Azure Key Vault), never in environment files checked into source control, and keep an audit log of rotation events separate from the secrets themselves — you'll need that trail if you're ever asked to prove when a compromised key stopped being accepted.

If you send webhooks too: SSRF is the other half of this problem

Everything above assumes you're receiving webhooks. If your product also sends webhooks to customer-supplied URLs — a common pattern for integration platforms and SaaS products with event notifications — you've built a system where your server makes outbound HTTP requests to addresses a user controls. That's a textbook SSRF primitive, and it's a live category of CVEs, not a theoretical one: multiple recent disclosures involve webhook dispatchers happily calling http://169.254.169.254/latest/meta-data/ (the AWS/GCP instance metadata endpoint) or internal RFC 1918 addresses because nothing validated the destination.

Registration-time URL validation isn't sufficient on its own, because DNS is mutable — a hostname that resolves to a public IP at registration time can be repointed to an internal address before the first delivery, or via a redirect on delivery. The defense needs two layers:

  • Registration-time check: reject non-HTTPS URLs (or require explicit opt-in for HTTP in development), reject anything that resolves to a private, loopback, or link-local range.
  • Delivery-time, resolved-IP check on every request: resolve the hostname, validate the IP — not the hostname string — against the blocked ranges (RFC 1918, 127.0.0.0/8, 169.254.0.0/16, multicast), and re-run that same check after following any redirect, since a 302 can point anywhere.

The most robust version of this runs the delivery workers on a network path with no route to your internal infrastructure at all — a dedicated egress proxy or isolated subnet — so a validation bug isn't the only thing standing between a malicious URL and your metadata endpoint.

A decision framework: build vs. adopt Standard Webhooks

If you're building a webhook sender from scratch today, the honest question is whether to hand-roll this signing and delivery layer or converge on the Standard Webhooks spec (or a provider built on it, like Svix). The spec already defines the headers (webhook-id, webhook-timestamp, webhook-signature), the exact signed-content format (id.timestamp.payload), and support for key rotation and Ed25519 signatures for consumers who want asymmetric verification instead of a shared secret.

ConsiderationHand-rollStandard Webhooks / managed provider
Time to shipDays, but security review adds weeksHours — spec and client libraries exist
Retry/backoff/DLQYou build and operate itIncluded if using a managed provider
Consumer integration costEvery customer writes custom verificationConsumers reuse existing spec-compliant libraries
Rotation supportYou design the overlap logicBuilt into the spec
Audit trail for complianceYou build delivery loggingTypically included

Hand-rolling still makes sense if you have exactly one or two consumers you control end-to-end and the integration surface is small enough that a bespoke scheme is genuinely cheaper to maintain than adopting a spec. Past that, the case for Standard Webhooks is less about the crypto — HMAC-SHA256 is HMAC-SHA256 either way — and more about not asking every integration partner to write bespoke verification code for your particular header names.

Failure modes and mitigations

  • Signature check passes but on a re-serialized body. Symptom: intermittent verification failures that seem to correlate with a specific client library or proxy. Fix: verify against raw bytes captured before any JSON parsing middleware runs.
  • Clock skew rejects legitimate webhooks. Symptom: valid deliveries failing the timestamp check under normal load. Fix: NTP-sync your servers, and don't set tolerance below ~60 seconds even if you're tempted to for security theater — real clock drift exists.
  • Idempotency table grows unbounded. Symptom: dedup lookups get slower over months. Fix: TTL cleanup job or a Redis-backed store with EX set at insert time instead of a manual sweep.
  • Rotation breaks delivery for slow-to-update consumers. Symptom: signature failures spike right after a scheduled rotation. Fix: dual-secret overlap window of at least 24 hours, communicated to integration partners in advance.
  • SSRF check only runs at registration. Symptom: a security audit or CVE report shows internal requests reaching a customer-controlled URL that passed registration but was later repointed via DNS. Fix: resolved-IP validation on every delivery attempt, not just at signup.

Working on this?

Syslabs' engineering team does architecture reviews on this kind of problem — webhook security, retry and delivery design, and the broader API development and integration and custom software development work it sits inside. Teams with compliance exposure around key management and access control also lean on our cybersecurity consulting practice. Book a 30-minute architecture review.