CrawlStateDAY 001 · 06 SEP 2026

The web is closing. Somebody has to write down when.

On 6 September 2026 we swept 100,000 domains and recorded what each one says about AI crawlers. We will do it again tomorrow, and every day after that, and keep all of it. This record is one day old. That is the point: in a year it will be a year old, and it cannot be started retroactively.

How the sweep works

Domains swept
100,000
Declared nothing
60,623
Priced their content
441
Census — 100,000 domains by rank, one sweep06 Sep 2026
rank 1 names an AI crawler    licence or price    content-signalrank 100,000

Four states. One is nearly all of it.

StateInkWhat it meansDomainsShare
BLOCKThe site names an AI crawler in its rules.12,61117.0%
CHARGEA licence is published, or a price is returned.4410.6%
SIGNALA Content-Signal preference, but no rule and no price.3760.5%
Nothing declared. Permitted by default, and we say so rather than guess.60,62381.9%

25,949 of the 100,000 did not answer at all and are excluded from the shares above. The blank row is not missing data. It is the finding.


The record — one column a dayDay 001
Reserved. 06 Sep 2026 – 05 Sep 2027
06 SEP 202605 SEP 2027

One column a day. Today there is one. The rest of the measure is reserved, not empty, and it is the only part of this product a competitor cannot buy, build, or catch up on.


Day 001, as recorded.

  • medium.comchargelicence declaredrsl
  • theguardian.comchargelicence declaredrsl
  • bild.dechargelicence declaredrsl
  • seattletimes.comchargelicence declaredrsl
  • drugs.comchargelicence declaredrsl
  • bustle.comchargelicence declaredrsl
  • telegraph.co.ukchargeHTTP 402402
  • dokuwiki.orgchargeHTTP 402402
  • nytimes.comnamednames 5 AI crawlersrobots
  • reddit.comundeclaredno AI directives
  • shopify.comundeclaredno AI directives
  • wikipedia.orgundeclaredno AI directives

Method.

Population
Tranco top 100,000, fetched 06 Sep 2026
Cadence
Once every 24 hours, unattended
Agent string
webpass-reality-check/0.1, with a contact address
Read
robots.txt, RSL licence links, Content-Signal, HTTP status
Spoofing
None. We never impersonate ClaudeBot or GPTBot.
Snapshot
snapshots/2026-09-06.json.gz · 100,000 records · 2.1 MB
Reachable
74,051 of 100,000
Recorded at
2026-09-06 15:15 UTC
Retention
Every snapshot, indefinitely

One limit, stated plainly: Cloudflare only returns a price to cryptographically verified crawlers. We are not one and we do not pretend to be, so a domain that charges privately reads here as undeclared. Every figure above is what we observed, not what we inferred.


What you can ask it.

Will this domain answer me?

GET /v1/domain/theguardian.com

{ "state": "charge",
  "since": "2026-09-06",
  "via":   "rsl" }

Which of my 50,000 will?

POST /v1/lookup

{ "allow":      31204,
  "block":       8113,
  "undeclared": 10683 }

What changed since Tuesday?

GET /v1/changes?since=…

{ "changed": 0,
  "note": "day 001 —
  no prior to diff" }

The third response is real. On day one there is nothing to compare against, and the API says so rather than returning an empty list.


15 SEP 2026Entered in advance

Cloudflare changes its defaults. Training and agent crawlers are blocked on ad-supported pages unless the owner opts out. Everything recorded above this line is the open web. Everything below it is whatever comes next.


Access.

PlanChargeIncludesSuits
Reading$01,000 lookups a monthTrying it
Indie$1925,000 lookups a monthOne person
Operator$199500,000 lookups a monthA crawler in production
The recordOn applicationUnlimited, plus the archiveYou want all of it

Open an account