The web is closing. Somebody has to write down when.
On 6 September 2026 we swept 100,000 domains and recorded what each one says about AI crawlers. We will do it again tomorrow, and every day after that, and keep all of it. This record is one day old. That is the point: in a year it will be a year old, and it cannot be started retroactively.
- Domains swept
- 100,000
- Declared nothing
- 60,623
- Priced their content
- 441
Four states. One is nearly all of it.
| State | Ink | What it means | Domains | Share |
|---|---|---|---|---|
| BLOCK | The site names an AI crawler in its rules. | 12,611 | 17.0% | |
| CHARGE | A licence is published, or a price is returned. | 441 | 0.6% | |
| SIGNAL | A Content-Signal preference, but no rule and no price. | 376 | 0.5% | |
| — | Nothing declared. Permitted by default, and we say so rather than guess. | 60,623 | 81.9% |
25,949 of the 100,000 did not answer at all and are excluded from the shares above. The blank row is not missing data. It is the finding.
One column a day. Today there is one. The rest of the measure is reserved, not empty, and it is the only part of this product a competitor cannot buy, build, or catch up on.
Day 001, as recorded.
- medium.comchargelicence declaredrsl
- theguardian.comchargelicence declaredrsl
- bild.dechargelicence declaredrsl
- seattletimes.comchargelicence declaredrsl
- drugs.comchargelicence declaredrsl
- bustle.comchargelicence declaredrsl
- telegraph.co.ukchargeHTTP 402402
- dokuwiki.orgchargeHTTP 402402
- nytimes.comnamednames 5 AI crawlersrobots
- reddit.comundeclaredno AI directives—
- shopify.comundeclaredno AI directives—
- wikipedia.orgundeclaredno AI directives—
Method.
- Population
- Tranco top 100,000, fetched 06 Sep 2026
- Cadence
- Once every 24 hours, unattended
- Agent string
- webpass-reality-check/0.1, with a contact address
- Read
- robots.txt, RSL licence links, Content-Signal, HTTP status
- Spoofing
- None. We never impersonate ClaudeBot or GPTBot.
- Snapshot
- snapshots/2026-09-06.json.gz · 100,000 records · 2.1 MB
- Reachable
- 74,051 of 100,000
- Recorded at
- 2026-09-06 15:15 UTC
- Retention
- Every snapshot, indefinitely
One limit, stated plainly: Cloudflare only returns a price to cryptographically verified crawlers. We are not one and we do not pretend to be, so a domain that charges privately reads here as undeclared. Every figure above is what we observed, not what we inferred.
What you can ask it.
Will this domain answer me?
GET /v1/domain/theguardian.com
{ "state": "charge",
"since": "2026-09-06",
"via": "rsl" }Which of my 50,000 will?
POST /v1/lookup
{ "allow": 31204,
"block": 8113,
"undeclared": 10683 }What changed since Tuesday?
GET /v1/changes?since=…
{ "changed": 0,
"note": "day 001 —
no prior to diff" }The third response is real. On day one there is nothing to compare against, and the API says so rather than returning an empty list.
Cloudflare changes its defaults. Training and agent crawlers are blocked on ad-supported pages unless the owner opts out. Everything recorded above this line is the open web. Everything below it is whatever comes next.
Access.
| Plan | Charge | Includes | Suits |
|---|---|---|---|
| Reading | $0 | 1,000 lookups a month | Trying it |
| Indie | $19 | 25,000 lookups a month | One person |
| Operator | $199 | 500,000 lookups a month | A crawler in production |
| The record | On application | Unlimited, plus the archive | You want all of it |