Skip to content

· skunxicat

On this page

From Inference to Proof: Web Bot Auth in the Wild

web-bot-auth

Bots are starting to sign their requests. Three things landed within days of each other: the Web Bot Auth working group adopted draft-ietf-webbotauth-httpsig-protocol as a working-group document (adoption, not approval — but a real signal the effort has momentum); curl 8.22.0 shipped with RFC 9421 HTTP Message Signatures support, putting the signing primitive into the client that half the internet’s automation shells out to; and a run of real signers — OpenAI, Cloudflare, DuckDuckGo among them — turned up signing requests to this very site. When the standards body, the ubiquitous client, and the biggest operators all move in the same week, it’s worth paying attention.

It’s tempting to file this under “another bot-detection signal,” a better rung on the ladder this log has been climbing. It’s bigger than that. Web Bot Auth is a small building block underneath a much larger shift: the web moving from preferences it can only hope bots respect to policy it can actually verify and enforce.

The problem: the web runs on preferences it can’t verify

For thirty years the contract between sites and bots has been robots.txt — a file that says “please don’t crawl here.” The word doing all the work is please. It’s a preference, not a control. A well-behaved crawler honors it; anything else ignores it, and the site has no way to know the difference at request time.

That was tolerable when automated traffic was a rounding error. It isn’t anymore. Bot traffic is not just growing, it’s the majority of what many sites serve, and the polite-preference model is visibly buckling under it. When a site decides it actually wants to apply a policy to automated traffic — rate-limit it, gate it, charge for it, serve it different content, or simply keep it out of the human-analytics numbers — it runs straight into two walls, and they’re mirror images of each other:

  • The declared bot can’t be verified. A request says it’s Googlebot. Maybe it is. Maybe it’s a scraper wearing the name. The user-agent is a string; claiming to be Google costs nothing.
  • The undeclared bot hides as a browser. Automation increasingly ships an ordinary Chrome/... Safari/537.36 user-agent — headless Chrome is now the crawler norm — so it doesn’t declare anything to verify in the first place. It just blends into human traffic.

Either way, a policy is only as good as the site’s ability to tell who it’s talking to. And today that telling is expensive and lives below the application layer: IP allow-lists that have to be maintained by hand and punish every future user of an address, PTR / reverse-DNS lookups, ASN cross-referencing — exactly the machinery this series built in the trust layer, and exactly the machinery that article admitted has a ceiling. All of it is the server reasoning about the client across the transport layer. None of it is the client telling the server who it is in a way the server can check.

That’s the gap. The earlier signals in this series each made spoofing harder without making verification cheap or certain:

  • JA4 (part one, part two) reads the TLS handshake instead of the user-agent — much harder to forge, but still a fingerprint, still probabilistic.
  • The trust layer cross-references IP against published CIDR ranges and ASN tiers — but it’s inference, and a real browser stack behind a residential proxy defeats it.

The draft’s own motivation names the same three broken identifiers from the inside: User-Agent (spoofable, and overloaded — an agent on Chromium wants to look like Chromium for rendering yet still distinguish its traffic); IP blocks (on cloud infra the address belongs to the platform, is re-published by the agent, and binds to nobody in particular — and dedicated blocks are expensive and carry reputation baggage); and shared secrets (a bearer token per site doesn’t scale past a few partnerships and rots as it spreads).

Web Bot Auth closes the gap with well-established cryptography. Instead of the server guessing, the bot proves it — a signature over the request, verifiable against a public key the bot itself publishes. It’s not “this fingerprint looks like Google.” It’s “this request carries a signature only Google’s private key could have produced, and here is Google’s public key to check it against.” Verification moves up to the request itself, and it’s cheap: fetch a key once, check a signature per request.

The identity is declared through Signature-Agent, then cryptographically verified. The user-agent no longer needs to act as the identity anchor.

How it works, briefly

Web Bot Auth is a profile of RFC 9421 HTTP Message Signatures. A signed request carries three headers:

  • Signature-Input — what was signed (which parts of the request), plus metadata: a keyid, an algorithm (ed25519 in practice), created/expires timestamps, and a tag="web-bot-auth" marking the profile.
  • Signature — the signature itself.
  • Signature-Agent — a URL identifying who signed it.

The keyid is a JWK thumbprint (RFC 7638) — a hash of the public key. To verify, you resolve the Signature-Agent URL to a set of keys, find the one whose thumbprint matches the keyid, and check the signature. There’s no single hard-coded place the keys live: the draft defines three discovery types, named by a type parameter on the Signature-Agent header —

  • directory (the default) — treat the value as an origin and fetch the key set from its well-known URI, /.well-known/http-message-signatures-directory.
  • jwks_uri — the value is a direct JWK Set URL.
  • cimd — the value is a Client ID Metadata Document URL, which in turn points to the keys.

The distinction isn’t bureaucratic. The well-known path is reserved, so directory proves something the others don’t: TLS authenticates the host, but nothing stops an arbitrary jwks_uri path from belonging to someone other than the host’s operator, whereas the well-known URI can only be served by whoever controls the origin. So directory binds the keys to a domain, not just a URL. Either way the shape is the same: no shared secrets, no pre-registration, no central authority. The bot publishes its keys; anyone can verify.

One caveat the draft is careful about: an unresolved Signature-Agent is only a claim. It becomes an identifier once its resolved key verifies the signature — and even then it establishes who signed, not that they are benevolent, authorized, or necessarily tied to a real-world operator. Identity, not permission.

What a site actually wants this for

“Catch spoofed bots” is the version of the problem this series started with, but it’s narrower than the real motivation. A companion draft, Nottingham’s Use Cases for Authentication of Web Bots, lays out what sites are actually trying to do, and most of it isn’t about catching liars at all:

  • Mitigating volumetric abuse. Today the blunt instrument is blocking by IP — which punishes every future user of that address and gives the wrongly-blocked no recourse. A stable identity that isn’t an IP lets a site rate-limit a bot, not an address.
  • Controlling access. Allow lists, deny lists, or conditional access (“you may crawl if you’re a non-profit / have paid / are certified”). Today this is robots.txt plus IP blocking — and robots.txt can ask but can’t enforce.
  • Serving different content. Some sites want to give a trusted bot more (or a distrusted one less — e.g. hiding commercially sensitive prices). That needs to know which bot, reliably.
  • Auditing behaviour. Verifying a bot actually honours the robots.txt it claims to, or getting trustworthy per-bot request metrics.
  • Classifying traffic. The one this log knows best: to understand your human audience you must exclude bots from analytics, and headless-browser crawlers make that harder every year. A signature is a clean “this one’s a bot” that no heuristic can match.
  • Authenticating site services. Health checks, CDNs talking to origins, uptime monitors — these already need to identify themselves, and today do it with IP allow lists and “magic” headers. Cryptographic identity is a straight upgrade.

Notice how many of these are problems this series kept hitting from the other side. The trust layer exists because IP and user-agent couldn’t reliably answer “which bot is this”; the use-cases draft is the same gap described by the people who operate the sites. Web Bot Auth is the identity primitive those use cases were missing.

And identity is only the floor. Once a request carries a verifiable signature, the same mechanism could carry more than who — it could carry why. The protocol lets an agent include an extra header expressing its intent and sign it too (§5.2.4), and the use-cases draft’s §3.4 argues for a standard, machine-readable way for a bot to convey authenticated context about its operation. What that context is — a stated purpose, “crawler versus user-triggered fetch,” an expected request rate — isn’t specified yet; no vocabulary for it exists in the drafts. But the shape of the shift is already visible: today a site infers intent from behaviour after the fact; a signed intent header would let a bot declare it up front, bound to a verified identity, so a site’s policy could act on it — “indexing crawlers may, training crawlers may not,” enforced instead of hoped. That’s the part that makes robots.txt look like the first draft of an idea rather than the finished one: a preference expressed by the site becomes a negotiation where the bot also makes verifiable claims. The vocabulary to express them is future work, not settled fact.

Catching it in the wild

The JA4 pipeline already captures telemetry at the edge: a CloudFront Function on the viewer-request event logs a line per request into the analytics pipeline. Web Bot Auth headers are ordinary request headers, so capturing them was a small extension of that same function — read the three headers, add them to the telemetry line, let them flow through the existing SNS → Firehose → S3 → Athena path into the same ja4_events stream.

No new infrastructure. The signature headers land right next to the JA4 fingerprint, the ASN, and the user-agent for the same request.

Then real signers showed up. Over a single 12-hour window, five distinct ones signed requests to this site:

SignerSignature-AgentCoversValidityUser-Agent
OpenAIchatgpt.com@authority @method signature-agent~1 hourordinary Chrome
Cloudflare Radarweb-bot-auth-directory.radar.cloudflare.com@authority signature-agent~5 minordinary Chrome
Cloudflare Workers (browser rendering)…cloudflare-browser-rendering-085.workers.dev@authority signature-agent~5 minHeadlessChrome
DuckDuckGo (DuckAssistBot)assistbot.duckduckgo.com@authority signature-agent~10 minDuckAssistBot/1.2 (self-declared)

All ed25519, all real operators sending live traffic. (A fifth signer showed up too — the REDbot diagnostic tool — but it’s a checker, not a bot operator, so it belongs to the negotiation story further down rather than this roster.) A few things jump out of that table.

Three of the four send an ordinary or headless browser user-agent — Mozilla/5.0 … Chrome/… Safari/537.36, and in the Cloudflare Workers case literally HeadlessChrome. Every classifier in the earlier articles, the CASE expressions and the JA4 lookups, would have filed these as plain desktop browsers. Nothing about the user-agent, and nothing about the TLS fingerprint, says “this is OpenAI.” The signature is the only thing that reveals them. That is the “beyond the user-agent” thesis reaching its logical end: the verified signature binds the request to a key published by the resolved Signature-Agent identity — something a scraper cannot reproduce merely by copying the headers, because it does not hold the private key.

DuckAssistBot is the interesting counter-case. It doesn’t hide — it declares DuckAssistBot/1.2 right in the user-agent, the old honest way — and it also signs. Here the signature isn’t revealing a concealed identity; it’s letting a site trust the declared one instead of taking it on faith. Same primitive, opposite motivation: one bot uses it to stop hiding, another to prove it was never lying.

And it showed up from a spray of different Azure IP addresses (ASN 8075) — every one carrying the same keyid. That’s the decoupling the use-cases draft promised made concrete: one cryptographic identity, many addresses, no allow-list to maintain. The reputation attaches to the key, not the infrastructure it happens to run on today.

A couple of wire-level details worth noting, because they’re what the traffic actually looks like rather than what the spec requires. Every one of these signers uses the legacy bare-string form of Signature-Agent ("https://chatgpt.com"), even though the -00 draft says signers MUST send the newer dictionary form and verifiers only MAY accept the bare string during migration — the deployed code is already behind the spec it implements. And validity windows range from OpenAI’s hour down to the five-minute posture the others take; the OpenAI and Cloudflare requests also carried a nonce for replay protection. These are the things you only learn by watching real traffic against a -00 draft while the ecosystem is still settling.

Proving it’s real

Capturing a Signature header proves nothing on its own — anyone can send bytes that look like a signature. The point of the whole scheme is that you can check it. So I ran all four through verification.

The verification is: take the keyid and the Signature-Agent URL, resolve it (every one used the default directory type, so that means fetching /.well-known/http-message-signatures-directory at that origin), find the published key whose RFC 7638 thumbprint equals the keyid, reconstruct the RFC 9421 signature base, and verify the ed25519 signature against that key.

Cloudflare has open-sourced web-bot-auth — Node libraries that handle the fiddly parts (signature-base construction, thumbprints, the crypto). So none of that had to be reimplemented; the only bespoke logic is the resolver that fetches a directory and matches the key. As a sanity check, computing the thumbprint of the key in Cloudflare’s own live reference directory reproduces exactly the kid they advertise — so the thumbprint implementation agrees with theirs.

This is full signature verification, not just a key-ID match: the library reconstructs the RFC 9421 signature base from the exact covered components each signer named (@authority, signature-agent, and for OpenAI also @method) and checks the ed25519 signature over it. To confirm the base was really being checked and not waved through, I flipped a covered component — changed the @authority — and verification failed, as it must. Only the untouched, correctly-reconstructed request passes.

The result: all four verify. OpenAI (chatgpt.com), Cloudflare (both the Radar directory and the Workers browser-rendering service), and DuckDuckGo (assistbot.duckduckgo.com) each check out — every signature’s full covered-component base validates against the key published at that signer’s own directory.

This is real. Not “a header that says OpenAI” — a signature whose full covered-component base checks out against OpenAI’s published key, and that breaks the moment any signed part of the request is altered. This is the first time in this entire series that the answer to “is this bot who it says it is?” is a yes backed by cryptography rather than a probably backed by a heuristic.

Asking a bot to sign: RFC 9421 §5

RFC 9421 has a second half that’s easy to miss. Section 5 defines how a server can request a signature it didn’t get, using an Accept-Signature response header — the same shape as Accept or Accept-Language, but for “please sign your next request, and here’s what I want covered.”

As a probe, this site now answers /signed-only that way. An unsigned request gets a 403 with:

Accept-Signature: sig=("@authority" "signature-agent";key="sig");created;expires;tag="web-bot-auth"
Cache-Control: no-store

A request that carries a signature gets a 200. (For now this only observes — the signature isn’t verified at the edge, because verifying it requires ed25519 and a network fetch that a CloudFront Function can’t do. That’s a deliberate line: capture and negotiate at the edge, verify off to the side.)

The results so far are split, and both are interesting. Most signers we saw — OpenAI, Cloudflare, DuckDuckGo — sign proactively: they attach a signature whether or not you ask, so they never see the challenge. But one client is built around the negotiation itself. The public REDbot checker signs challenge-driven, not proactively — its recently merged Web Bot Auth support sends requests unsigned by default and signs only when the origin answers with a 401/403/429 carrying Accept-Signature, then transparently retries the same request signed. Against /signed-only that’s exactly what happened: unsigned request, 403 challenge, signed retry — a real, in-the-wild client completing the §5 round-trip end to end.

One honest footnote on that: REDbot deliberately ignores the specific components an Accept-Signature asks for and sends its standard signature — the reasoning (sound, per the draft) being that a challenger must request the same parameters the signer already produces; echoing a server-chosen component set would be later work if it’s ever needed. Fine print of a protocol still at -00, whose own reference clients are actively changing.

Server-requested signing is otherwise barely deployed. But the mechanism works end to end, and it’s the forward-looking half of the standard: a day may come when a site can say “prove who you are or you don’t get in,” and the bots that matter can answer.

Where this sits

It’s worth being precise about what this is and isn’t, because the earlier articles earned their credibility by not overclaiming and this one shouldn’t spend it.

What it is: the first bot-identity signal that’s a cryptographic fact instead of an inference. When a request verifies, you know who sent it, to the same degree you trust that OpenAI controls chatgpt.com and keeps its signing key private.

What it isn’t: widespread (a handful of signers so far, all of them large, well-resourced operators), finished (-00, freshly adopted, the wire format still shifting), or authorization (a verified signature tells you who, not what they’re allowed to do — that’s your policy, not the protocol’s). And it’s identity for the well-behaved: a bot that wants to be recognized will sign; a bot that wants to hide simply won’t, and for those, the JA4 and trust layers are still the only game. Web Bot Auth doesn’t replace fingerprinting. It adds a lane for the bots that choose to identify themselves — and, increasingly, the big ones are choosing to.

And it isn’t uncontroversial. The same use-cases draft is candid that the working group doesn’t agree on scope, and the disagreements are real ones. Standardising bot authentication could nudge the web toward a place where every bot is expected to authenticate — raising the barrier to launching a new one, and pressuring toward centralisation because the incumbents are the ones who can absorb that cost. It shifts the balance of power between sites and bots: welcome if you think sites are drowning in an AI-crawler onslaught, worrying if you think it hands powerful sites a finer-grained lever to decide who gets to read them — including accessibility tools and user-driven agents. It was even argued at IETF 124 that the resource requirements would disproportionately favour already well-resourced players. None of this is settled, and this post isn’t taking a side on it — but a signal this powerful arriving this fast deserves the caveat stated plainly, not buried.

The arc of the series, then, in one line: from fingerprints (harder to fake) to fingerprints plus reputation (the trust layer) to cryptographic proof (web bot auth). Each rung is a stronger answer to is this bot who it says it is? — and this is the first rung where the answer isn’t a guess.

But the reason it matters isn’t really the answer to that one question. It’s that a verifiable answer is the thing every policy was missing. robots.txt asked bots to behave and hoped; verification, when it existed at all, meant expensive IP and DNS gymnastics below the application layer. Web Bot Auth puts a cheap, checkable identity — and, increasingly, a declared intent — into the request itself. That’s the building block a web of enforceable, negotiated policy between sites and bots gets built on. Four real operators signing in a single day — OpenAI, two Cloudflare services, DuckDuckGo — plus a diagnostic tool exercising the challenge-response path, is a small beginning. The shift underneath it isn’t.


Part of the bot-detection series: Beyond the User-Agent: Catching Spoofed Bots with JA4 and ASN, Building a JA4 Fingerprint Pipeline on AWS, and Building a Trust Layer.