Perspectives

The web is splitting into three tiers: open, blocked, and paid. Most brands haven't decided which one they are in.

GE

Gensight.AI

June 4, 2026

The web is splitting into three tiers: open, blocked, and paid. Most brands haven't decided which one they are in.

For about twenty years, the implicit deal of the open web was simple. You published content, search engines indexed it, people found it, and you monetised their attention through ads, subscriptions, or the eventual sale. The reader was a human being with eyes, and the entire economic model assumed those eyes would arrive.

That assumption is breaking. When an AI agent reads your page, extracts the answer, and completes the task on the user's behalf, the human never arrives. They do not see your ad. They do not click your affiliate link. They do not land on the page where the conversion happens. Your content did the work, and the value was captured somewhere you cannot monetise. Industry data through 2025 and into 2026 shows AI-origin traffic rising sharply while the referral traffic AI sends back remains a fraction of what it takes, with the publishing sector absorbing the majority of the crawl load.

This has produced a real question that did not exist three years ago: who is your content actually for? And the infrastructure layer has already started answering it on your behalf.

The default flipped, and most people did not notice

On 1 July 2025, Cloudflare, which sits in front of roughly a fifth of the web, made blocking AI crawlers the default for every new domain. The model moved from opt-out to opt-in: AI companies now need explicit permission to crawl a Cloudflare-protected site that has not granted it. Over two and a half million sites have moved to disallow AI training access. Cloudflare also introduced pay-per-crawl, where a blocked crawler receives an HTTP 402 Payment Required response and the publisher can set a price for access.

The result is that the web is sorting itself into three tiers. Open, where crawlers pass freely. Blocked, where they are refused. And paid, a genuinely new category, where access has a price. This is the most significant change to how machines access web content since robots.txt, and it is happening largely through a default setting that most site owners never consciously chose.

The mistake: treating this as one decision when it is three

Here is where the decision is quietly being made for people, rather than by them. The block-or-allow choice is being treated as a single binary, when the infrastructure already distinguishes three different reasons a machine might visit. Cloudflare's own crawler controls separate them explicitly: crawling to train a model, crawling for inference, and crawling for search.

Those are not the same thing, and conflating them is the error.

Blocking training crawlers is defensible and often wise. If a model absorbs your content into its weights with no attribution and no payment, you have handed over an asset for nothing, and gating or charging for that access is a reasonable commercial decision. This is the tier where the licensing and pay-per-crawl conversation properly belongs.

Blocking search and inference crawlers is usually self-sabotage. This is the layer that surfaces you when someone asks ChatGPT, Perplexity, or Claude for a recommendation. A brand that flips the default to block-all, whether to save on bandwidth or to protest training use, can quietly make itself invisible to the exact retrieval systems that are becoming the new front door. They think they are protecting their content. They are actually removing themselves from the shelf.

The brands getting this right are separating the two decisions: gate or charge for training access, stay open and well-structured for retrieval. The risk is making one blunt choice with a toggle you do not fully understand, and never realising the two were different decisions at all.

Why this is the real GEO question now

For the past year, the discipline around AI visibility has mostly asked: are you optimised? Is your schema clean, your entity resolution tight, your content retrievable? Those questions still matter. But they sit underneath a more fundamental one that the closing web has surfaced: have you decided what you are optimising for, and to whom you are visible?

Visibility is becoming a deliberate access-tiering decision. A brand now has to answer, per crawler type, what it wants: to be trained on or not, to be retrievable or not, to be free or paid. The optimisation work is downstream of that decision. There is no point investing in perfect retrieval readiness if a default setting three layers up is returning a 402 to the very crawlers you need.

This also reframes the creator-economy anxiety. The fear is that agents will hollow out the attention economy, and for pure ad-funded models that pressure is real. But the response is not to wall everything off. It is to recognise that content now serves two distinct audiences with two distinct monetisation paths. Humans, monetised through attention. Machines, monetised through being the cited, recommended source, and increasingly through licensing the training access directly. Optimising for one is not optimising for the other, and the access tier is how a brand chooses its mix.

What to actually do

Three practical steps follow, and none of them is "block everything" or "open everything."

First, audit what your current setup is actually doing. A surprising number of brands do not know whether their infrastructure is returning a 402 or a 200 to AI search crawlers, because the default was set for them. The single most common own-goal right now is a brand that wants to be recommended by AI, while unknowingly blocking the crawlers that would make that possible.

Second, separate your training policy from your retrieval policy. Decide deliberately whether you want your content used for model training, and whether you want to charge for it. Then, separately, ensure the search and inference crawlers that drive recommendation are allowed and well served. These are different levers and they should be set independently.

Third, treat retrievability as a commercial asset, not a technical afterthought. In a world where the reader is increasingly a machine acting for a human, being easy to retrieve, parse, and cite is how you stay on the shelf. The brands that win the next phase will be the ones that made the access-tier decision on purpose, rather than inheriting it from a default.

The open web is not ending. It is being priced, tiered, and gated, deliberately, for the first time. The question is no longer just whether you are visible to AI. It is whether you have decided, on purpose, which machines you want to see you, for what, and at what price.

Ready to stop monitoring and start dominating?

Run Your Free Baseline Audit