The crawler tolls, one year on
Default blocking and pay-per-crawl changed who can read the web. An assessment of what actually happened.
A year ago today a CDN sitting in front of a large fraction of the web flipped its default: AI crawlers blocked unless explicitly allowed, with a marketplace for charging per crawl.
Enough time has passed to say what actually happened rather than what everyone predicted.
what changed#
The norm inverted. Before, crawling was permitted by default and robots.txt was a request. Now, for a large share of the web, crawling is denied by default and access is a negotiation.
That is a genuine structural change to how the web works and it happened through one company's configuration default rather than through any standards process, legislation, or public deliberation.
Licensing deals concentrated. Large AI companies negotiated bulk access with large publishers. That was always the likely outcome: the parties with lawyers and leverage made arrangements, and the arrangements are private.
Small publishers got very little. The pay-per-crawl mechanism works, technically. The revenue for a site with modest traffic is negligible — the arithmetic never supported anything else. The publishers who most needed a new economic model got the one that pays least.
Non-commercial crawling got harder. Academic researchers, archivists, and independent search projects have no licensing budget and no negotiating position. Carve-outs exist and are discretionary, which means the ability to study the web is now something you apply for.
That is the outcome I was most worried about and it is the one that materialized most clearly.
what did not change#
Training data supply. The frontier labs have enormous existing corpora, licensed sources, and synthetic generation. The marginal value of newly crawled web text was already declining. Restricting it did not create the leverage publishers hoped for.
Traffic. Referral traffic from search to publishers continued its decline, driven by AI answers in search results, which is a completely separate mechanism from training crawlers and which blocking crawlers does nothing about.
This is the part that was most misunderstood at the time. The traffic problem and the training problem have different causes and blocking crawlers only addresses one of them — the one with less economic impact.
the thing to actually take from it#
Infrastructure defaults are policy. A configuration change at a chokepoint reshaped access to a large fraction of the web, with no process and no appeal.
That is not a criticism of the specific decision, which was popular and defensible. It is an observation about where power actually sits, and it generalizes: the entities that can change the web's behavior are the ones with concentration at a layer everyone depends on, and there are about five of them.
For your own site, the decision remains yours and it is worth making deliberately rather than accepting a default:
- Documentation sites frequently want to be in the training data. Being the thing the model knows about is worth more than the pageview you did not get.
- Original reporting and analysis has a stronger case for restriction.
- Anything you want found should still permit search crawlers, which are a different category and are frequently blocked by accident when people configure this.
Check what you are actually blocking. A meaningful number of sites blocked their own search indexing in the first months of this and did not notice for weeks.
the unresolved thing#
The web's economic model — publish freely, get traffic, monetize traffic — is breaking, and nothing has replaced it.
Crawler tolls are not the replacement; the arithmetic does not work at the scale of the actual web, where most content is made by people with no ability to negotiate anything.
Licensing deals are not the replacement either; they work for a few hundred large publishers and for nobody else.
I do not know what the replacement is. I am increasingly convinced that nobody does, and that the interval between the old model failing and a new one existing is going to be long and is going to be bad for the open web.
That is a genuinely pessimistic conclusion and I have not found a way around it in a year of thinking about it.
— Dom, July 1, 2026