Webhook Security Is Distributed Systems Security

Webhook Security Is Distributed Systems Security

A webhook is not an HTTP request. It is an untrusted, replayable, out-of-order, at-least-once message stream that happens to arrive over POST, and it has a public address. Most webhook bugs are not injection. They are the distributed-systems failures you already know (replay, duplicate delivery, reordering, retry storms) showing up at an endpoint nobody treated like a message consumer. Here is the trust boundary, the five failures, and the actual checks, at a depth you can run against your own endpoint.

October 2, 2026
Harrison Guo
8 min read
Security Architecture

A webhook looks like the easiest thing in your system. A provider sends you a POST, you read the JSON, you do the thing. Three lines.

It is one of the most dangerous. Not because the parsing is hard, but because a webhook is not really an HTTP request. It is a message from a queue you do not own, with no delivery guarantees, arriving at an address anyone on the internet can find. At-least-once, unordered, retried, and signed by a party outside your trust boundary.

Almost every webhook bug I have seen is a distributed-systems failure wearing an HTTP costume. Replay. Duplicate delivery. Reordering. Retry storms. You already know these failures from building queue consumers. The trap is that the webhook arrives over POST, so nobody treats the handler like the message consumer it is.

Here is the trust boundary, the five failures, and the checks that actually hold. All of it runs against an endpoint you control.

The signature is the entire trust boundary

The endpoint is public. The only thing separating a real event from one I forged in curl is the signature. So the signature check is not a formality. It is the whole wall.

The bug that defeats it is almost always the same, and it is not a weak algorithm. It is verifying the wrong bytes.

The sender computes an HMAC over the exact body it transmitted. Your web framework then helpfully parses that body into an object. If you verify the signature over a re-serialized version of that object, you are hashing different bytes. Key order changed, whitespace changed, a number got reformatted. The verification fails on legitimate events. The usual fix a tired engineer reaches for is to make it pass, which quietly means not really verifying at all.

Verify the raw body, before anything touches it.

import hmac, hashlib

def verify(raw_body: bytes, header_sig: str, secret: bytes) -> bool:
    expected = hmac.new(secret, raw_body, hashlib.sha256).hexdigest()
    # constant-time: never use == on a secret-derived value
    return hmac.compare_digest(expected, header_sig)

Two things people get wrong even here. They pass the parsed-and-re-dumped body instead of raw_body. And they compare with ==, which short-circuits on the first differing byte and leaks the correct prefix through timing. compare_digest exists for exactly this. On your own endpoint, send a valid event, then flip one byte of the body and confirm it is rejected. Then re-dump the parsed JSON and verify against that, and watch a real event get rejected. That is the bug, reproduced.

A valid signature, replayed, is still valid

Here is the step that signature-only designs miss. I capture one real, correctly-signed delivery. I send the exact same bytes again tomorrow. The signature still verifies, because nothing about it changed. If that second delivery credits an account or ships an order, I just did it twice with a replay.

A signature proves origin and integrity. It says nothing about freshness. Freshness needs a timestamp inside the signed payload and a window.

import time

def fresh(signed_timestamp: int, now: int, window=300, skew=30) -> bool:
    # the timestamp must be INSIDE the signed region, or an attacker edits it
    if signed_timestamp > now + skew:
        return False            # reject the future outright; it buys replay runway
    return now - signed_timestamp <= window

The timestamp has to be part of what the HMAC covers. This is why good providers sign t=<timestamp>.<body>, not just the body. If the timestamp sits in an unsigned header, I change it on replay and the window is useless. Five minutes is a common window. Pair it with rejecting a timestamp from the future, which otherwise lets an attacker buy an arbitrarily long replay runway.

A window narrows the replay window. It does not close it, because within those five minutes a replay is still valid. Closing it needs the next failure’s fix.

At-least-once means you will see duplicates in normal operation

This one is not even an attack. It is Tuesday.

Webhook delivery is at-least-once. The provider sends the event, your server is slow to ack or a connection blips, the provider does not get its 2xx in time, so it sends again. Same event, delivered twice, both legitimately signed. If your handler is not idempotent, “we got charged twice” and “the order shipped twice” are now your bug, caused by the network doing exactly what at-least-once delivery promises.

The fix is a dedup key on the event’s own id, checked before you act, stored atomically.

def handle(event):
    # INSERT ... ON CONFLICT DO NOTHING, or a unique constraint on event_id.
    # The check and the claim must be one atomic operation, or two concurrent
    # deliveries both pass the check before either writes.
    if not claim_event(event["id"]):   # returns False if already seen
        return 200                     # ack and drop: it is a duplicate
    do_the_work(event)
    return 200

The subtle part is atomicity. A read-then-write (“have I seen this id? no, ok process it”) loses to two deliveries arriving at once: both read “no”, both process. The dedup has to be a single atomic claim, a unique constraint or an ON CONFLICT, not an if followed by an insert. This is the same hazard as any exactly-once-effect consumer, which is why it belongs in the idempotency conversation, not a webhook-specific one.

Dedup also finally closes the replay hole from the previous section. A replayed event carries the same id, so an idempotent handler drops it whether the replay is malicious or just the provider retrying.

Order is not guaranteed, and your state machine assumes it is

Two events, subscription.created and subscription.cancelled, sent a second apart. They can arrive in either order. Retries make it worse: the first delivery fails, gets retried, and lands after an event that was sent later. Now your handler processes cancelled then created, and you have a live subscription that the provider’s records say is dead.

You cannot fix ordering at the transport. You fix it in the handler by not depending on arrival order.

  • Carry the entity’s version or a monotonic sequence in the event, and ignore an event older than the state you already hold.
  • Make the effect a function of the final state, not of the transition, where you can. Setting status to a value is order-independent in a way that incrementing a counter is not.
  • When you truly need ordering, the event only tells you something changed. Go read the current truth from the provider’s API and act on that.

This is the same reasoning as choosing who owns completion between RPC and NATS. The moment delivery is asynchronous, arrival order stops being a thing you can lean on, and the consumer has to be written to survive it.

Retries become a storm, and the storm can be aimed

Non-2xx means retry. So if your handler does its slow work and then returns a 500 because the work failed, the provider retries, your handler does the slow work again, fails again, and you have built a feedback loop that multiplies load exactly when you are already degraded. A provider with aggressive retries can turn one bad deploy into a self-inflicted flood.

Separate acknowledgement from processing. Verify the signature, claim the id, write the raw event to a queue of your own, return 2xx. Process off that queue.

def endpoint(request):
    if not verify(request.raw_body, request.headers["X-Signature"], SECRET):
        return 401                      # do not retry an unauthenticated caller
    if not fresh(signed_ts(request), int(time.time())):
        return 401
    enqueue(request.raw_body)           # durable, your side
    return 200                          # ack fast; processing happens later

Now a processing failure is your queue’s retry policy, with your backoff and your dead-letter, not the provider hammering your front door. A poison event that can never succeed lands in a dead-letter queue for a human, instead of retrying until the provider gives up and you silently lose it. Return non-2xx only for “I could not authenticate or accept this”, never for “I accepted it but the work failed”.

The endpoint is attack surface, in both directions

Two more, quickly, because they are the ones a checklist catches and a design review misses.

Inbound: the URL is public and guessable. Rate-limit it. An unauthenticated flood of garbage that each triggers a full signature computation is a cheap denial-of-service if verification is your most expensive early step.

Outbound: if your product lets a user configure where you send webhooks, that is server-side request forgery waiting to happen. A user points their webhook URL at http://169.254.169.254/ or an internal admin service, and your server dutifully makes the request from inside your network. Validate the destination against a deny-list of private ranges and the cloud metadata address, and re-run that check on every redirect hop. A hostname deny-list alone is not enough: a name that passes your check can re-resolve to 169.254.169.254 by the time you connect (DNS rebinding), so validate the IP you actually connect to and pin it, not just the hostname. Treat the whole thing as the SSRF primitive it is.

The one sentence to take

Stop reading a webhook as an HTTP request and start reading it as a message from a queue you do not own. Every hard property follows from that: the producer is untrusted so you verify the raw bytes, delivery is at-least-once so you dedup, order is not guaranteed so you version, retries are automatic so you ack fast and process async, and the address is public so it is attack surface.

None of this is exotic. If you have built a serious message consumer, you have solved all of it before. The only new thing a webhook adds is that you cannot trust the sender, which raises the stakes on the exact same list. The engineers who get webhooks wrong are not missing a security trick. They are treating a distributed-systems problem as a parsing problem.

🎧 More Ways to Consume This Content

I occasionally advise small teams on backend reliability, Go performance, and production AI systems. Learn more: /services

Comments

This space is waiting for your voice.

Comments will be supported shortly. Stay connected for updates!

Preview of future curated comments

This section will display user comments from various platforms like X, Reddit, YouTube, and more. Comments will be curated for quality and relevance.