AI crawlers are consuming the web far faster than the AI products built on them send visitors back, and the websites that produce the best information are responding by shutting the door. That is the argument of a piece published on 31 August 2026 by Dana McKay and Damiano Spina of RMIT University in The Conversation, syndicated the same day by TechXplore. Their conclusion is uncomfortable for everyone: as reputable publishers block AI crawlers, AI answers lean harder on the sites that do not block, many of which are low quality or themselves AI-generated.
The mechanism is not complicated. For thirty years the web ran on a deal: search engines could crawl a site for free, and in return they sent readers to it. Generative AI tools now crawl to train models and to write answers, and the link to the source has become an optional extra. Ranking the web was a natural language processing problem long before chatbots existed; what has changed is that the answer now appears above the results, and the click that paid for the content does not happen.
This article sets out what the RMIT authors said, the traffic data behind it, why AI crawlers cost site owners money, how publishers and Cloudflare are blocking them, the peer-reviewed evidence that this quality trap is real, and what it means for anyone searching for information or running a website. For businesses that depend on search, it connects to the SEO services and GEO services conversations already under way on this site. Every figure is traceable to a source in the References section.
Table of contents
- What the RMIT Authors Say About AI Crawlers and the Web’s Social Contract
- How Much Traffic AI Answers Are Really Eating
- Why AI Crawlers Cost Website Owners Money
- How Websites Are Blocking AI Crawlers
- The Quality Trap: Reputable Sites Block AI Crawlers, Misinformation Sites Do Not
- What Pay-to-Crawl and Licensing Deals Would Change
- What It Means for You: Finding Reliable Information Around AI Crawlers
- What UK Businesses Should Do About AI Crawlers on Their Own Website
- Frequently Asked Questions
- References
What the RMIT Authors Say About AI Crawlers and the Web's Social Contract
McKay, an Associate Dean at RMIT’s School of Computing Technologies, and Spina, a Senior Lecturer in the same school and a member of the ARC Centre of Excellence for Automated Decision-Making and Society, frame the problem as the collapse of a social contract rather than a glitch in the AI tools themselves. The piece carries a DOI (10.64628/AA.rxcnppj3t) and discloses that McKay has received funding from Google, which makes its criticism of AI summaries more notable, not less.
The thirty-year deal between crawlers and websites
In the early web, the authors write, “search engines and content creators came to an agreement about crawling”. Sites were free for search engines to access, and search engines were allowed to reproduce small snippets. In return, “search engines provided links to the sites owned by content creators, who benefited from that web traffic”. Anyone who disliked the deal could opt out with a robots.txt file. The deal was informal, but it funded most of what people read online.
How AI crawlers broke the deal
The break, in their words, is that “AI tools are crawling sites not to link to them, but to train models and generate answers (which may or may not be accurate)”. ChatGPT and Google’s AI Overviews “may still include links to sources, but they’re a kind of optional extra to the main answer”. AI crawlers therefore take the content that made search useful while returning almost none of the traffic that paid for it. They also “crawl more deeply and more intensely than traditional web crawlers”, so each visit costs the site more.
The bad dynamic for everyone, including AI companies
The consequence the authors describe is a loop. Websites lose traffic and revenue, so they block AI crawlers. AI answers then “depend more on low-quality websites (many of which are also generated by AI)”. Good information becomes harder to find, which damages the AI companies’ own products. “As a result,” they conclude, “good information can be harder than ever to find.” The table below sets the old deal against the new one.
| Element of the deal | Search-engine era (1995 to 2023) | AI crawler era (2024 onwards) |
|---|---|---|
| What the crawler takes | An index entry and a short snippet | Full text for model training and answer generation |
| What the site gets back | A ranked link and referral visits | A citation that is “an optional extra”, if anything |
| Cost of being crawled | Low; crawlers were shallow and infrequent | Higher; AI crawlers crawl “more deeply and more intensely” |
| Opt-out mechanism | robots.txt, broadly honoured | robots.txt, which “some AI companies may ignore” |
| Who funds the content | Advertising and subscriptions from referred visitors | Unresolved; “pay to crawl” models have “not gained traction” |
| Effect on information quality | Sites competed to be worth linking to | Reputable sites withdraw; AI answers draw on what remains |
How Much Traffic AI Answers Are Really Eating
The authors point to a looming event they call “Google Zero”, a phrase coined by The Verge in 2024 for the day when referral traffic from Google drops to nothing. That day has not arrived, but the data on how AI summaries change behaviour is now solid, and it comes from independent measurement rather than publisher complaints.
Pew Research: clicks halve when an AI summary appears
The Pew Research Center tracked the browsing of more than 900 US adults during March 2025, covering 68,879 unique Google searches. When an AI summary appeared, users clicked a traditional result on 8 percent of visits; when it did not, they clicked on 15 percent. They clicked a link inside the summary itself on just 1 percent of visits, and they ended their browsing session entirely on 26 percent of pages with a summary against 16 percent without one.
Wikipedia lost around 15 percent of search traffic to AI Overviews
The RMIT piece cites a study of Wikipedia by Mehrzad Khosravi and Hema Yoganarasimhan, available on arXiv and SSRN, which used the staggered rollout of AI Overviews across countries and languages as a natural experiment. Across 161,382 matched article-language pairs, exposure to AI Overviews cut daily traffic to English articles by roughly 15 percent, with the largest declines in culture topics and smaller ones in STEM. Scaled to their 52,262-article sample, that is about 11.5 million fewer daily visits.
The Wikimedia Foundation confirmed the drop from its own logs
The Wikimedia Foundation’s own analysis, published on its Diff blog in October 2025, found human pageviews down about 8 percent between March and August 2025 compared with the same months a year earlier, after it reclassified a wave of bots built to evade detection that had been masquerading as human readers. The foundation attributed the decline to search engines answering questions directly, often using Wikipedia’s own content. The chart below gathers the Pew figures in one place.
Why AI Crawlers Cost Website Owners Money
Lost referrals are only half the problem. The other half is that AI crawlers generate real server load without the page views, ad impressions or subscriptions that human visits bring. The RMIT authors call this “a bad dynamic for website owners, the public, and even AI companies themselves”, and Cloudflare’s network data explains why.
Bots now outnumber humans on the web
On 5 June 2026 Cloudflare chief executive Matthew Prince reported that automated traffic had reached 57.3 percent of worldwide HTTP requests for HTML content, against 42.7 percent for humans, as covered by Search Engine Land. Prince had predicted the crossover for 2027; agentic browsing brought it forward by more than a year. The RMIT piece rounds this to “over half of all web traffic is now AI bots”, noting that Cloudflare manages 30 percent or more of the top 10,000 sites, so its Radar figures are the best public proxy for the whole web.
AI crawlers fetch far more than they refer
Cloudflare publishes a crawl-to-refer ratio on Radar for each AI platform: the number of pages its AI crawlers fetch for every human visitor the platform sends back. For search engines that ratio has historically been in single figures; for pure AI answer engines it runs into the hundreds or thousands. Prince’s own framing is that a shopper might visit five sites while an AI agent visits thousands, “creating real traffic and real server load without the same clicks, ad views, or customer relationships”. Cloudflare also told TechCrunch that more than 50 percent of AI crawler traffic is spent re-fetching pages that have not changed.
robots.txt is a request, not a wall
The web’s only universal opt-out is voluntary. The RMIT authors note that “some AI companies may ignore this polite request”, citing Reuters’ June 2024 report that several AI firms were bypassing the standard to scrape publisher sites. Digiday reported in June 2026 that a TollBit analysis found 30 percent of AI bot scrapes in the fourth quarter of 2025 ignored explicit robots.txt permissions. The chart below sets the Cloudflare and TollBit figures side by side.
How Websites Are Blocking AI Crawlers
Blocking has moved from a niche robots.txt edit to a default posture at some of the largest publishers, and from 15 September 2026 it becomes a default at the network level for many Cloudflare sites. The RMIT authors describe both developments; the primary sources add the detail.
Publishers switch from block-lists to allow-lists
Digiday reported on 9 June 2026 that Reuters and Time had both moved from allowing every bot by default to blocking all of them and maintaining an approved list. Josh London, head of Reuters Professional, said the company “went from a default allow-all to a default disallow all” because “our content costs money to create”. Time now admits about 70 bots, according to its chief operating officer Mark Howard. People Inc. found that switching from a block-list to an allow-list took it from blocking roughly 2,100 user agents to more than 30,000, a measure of how many AI crawlers now exist.
Cloudflare’s 15 September default
On 1 July 2026 Cloudflare announced in its “Your site, your rules” post that it would classify bots by three behaviours, Search, Agent and Training, and that from 15 September 2026 Training and Agent crawlers would be blocked by default on pages that display ads, while Search stays allowed. Multi-purpose crawlers that combine Search with Training, which Cloudflare names as Googlebot, Applebot and BingBot, will be judged by their most restrictive behaviour. TechCrunch reported that the new defaults apply to new customers, new sites from existing customers and all existing free-tier customers.
What that means for AI Overviews
Because Googlebot crawls for both Search and AI features, the RMIT authors reason that up to 30 percent of the world’s top sites, Cloudflare’s share, “will no longer appear in Google AI Overviews summaries”. That figure is their extrapolation rather than a Cloudflare statement, and site owners can opt out of the new defaults before 15 September. Even so, the direction is clear: the pages that carry advertising, which is to say the professionally produced ones, are the pages AI crawlers will increasingly be kept away from.
| Response | Who is doing it | How it works | Limits |
|---|---|---|---|
| robots.txt disallow rules | 60.0% of reputable news sites (Steinacker-Olsztyn et al.) | Names AI user agents that may not crawl | Voluntary; 30% of scrapes ignored it in Q4 2025 (TollBit) |
| Default-deny allow-lists | Reuters, Time, People Inc., The Atlantic | Block every bot, then approve about 70 (Time) | Needs a vendor such as ScalePost to manage |
| Active blocking by user agent | Both reputable and misinformation sites | Server refuses content to AI crawler user agents | Easily evaded by a bot that lies about its identity |
| Network-level defaults | Cloudflare, from 15 September 2026 | Training and Agent bots blocked on ad pages; Search allowed | New domains and free tier first; owners can opt out |
| Pay-per-crawl and pay-per-use | Cloudflare with Ceramic.ai and You.com | Publisher paid when content is fetched or used | “Haven’t gained traction” (RMIT authors) |
| Direct licensing deals | Reuters and other large publishers | Blocking creates the friction that brings AI firms to the table | Only available to publishers with scale |
| Cloudflare bot class | What it does on your site | Examples Cloudflare gives | Default on ad pages from 15 Sept 2026 |
|---|---|---|---|
| Search | Indexes content to answer queries later; expected to send referrals | Classic search indexing | Allowed |
| Agent | Acts in real time on a person’s behalf to complete a task | ChatGPT-User, browser-use agents driving Chrome | Blocked |
| Training | Absorbs content permanently into a model | Dedicated training crawlers | Blocked |
| Multi-purpose | Combines Search with Training or Agent use | Googlebot, Applebot, BingBot | Judged by the most restrictive applicable rule |
The Quality Trap: Reputable Sites Block AI Crawlers, Misinformation Sites Do Not
The most important claim in the RMIT piece is also the best evidenced: blocking is not evenly distributed. Credible publishers are far more likely to shut out AI crawlers than sites that spread misinformation, so the pool AI answers draw from tilts towards the unreliable. Three peer-reviewed studies from 2025 and 2026 make that case.
Sixty percent of reputable news sites block at least one AI crawler; nine percent of misinformation sites do
“Is Misinformation More Open? A Study of robots.txt Gatekeeping on the Web”, by Nicolas Steinacker-Olsztyn, Devashish Gosain and Ha Dao, was published in the Proceedings of the ACM Web Conference 2026 and is available on arXiv. Among sites with a robots.txt file, 60.0 percent of reputable news sites disallowed at least one AI crawler, against 9.1 percent of misinformation sites. Reputable sites forbade an average of 15.5 AI user agents; misinformation sites forbade fewer than one. The gap has widened: AI-blocking by reputable sites rose from 23 percent in September 2023 to nearly 60 percent by May 2025.
Sites that block AI crawlers are less likely to be cited, even when the content is reachable
A SIGIR 2026 paper by Grossman and colleagues, “How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews” (ACM), compared 11,500 real user queries across Google Search, AI Overviews and Gemini. AI Overviews appeared for 51.5 percent of queries. Traditional search was significantly more likely to surface government and education sites, while generative search favoured Google-owned content. Crucially, “websites that block Google’s AI crawler are significantly less likely to be retrieved by AIOs, despite having access to the content”, which is the retrieval-augmented generation loophole the RMIT authors describe.
One in six sources cited by AI search is itself AI-generated
Mowafak Allaham and Nicholas Diakopoulos audited ChatGPT, Copilot, Gemini and Perplexity with 712 real queries on politics, health and the environment (arXiv, May 2026). Around 16 percent of cited sources were AI-generated websites, across all four engines. The authors of the RMIT piece draw the obvious line to model collapse, the phenomenon documented in Nature in 2024, where models trained on recursively generated text degrade. If AI crawlers can only reach the sites that did not bother to block them, the training and grounding data both get worse.
| Study | What it measured | Headline finding |
|---|---|---|
| Pew Research Center, July 2025 | Browsing of 900+ US adults, 68,879 Google searches | Clicks on results fall from 15% to 8% when an AI summary appears |
| Khosravi and Yoganarasimhan, 2026 | 161,382 Wikipedia article-language pairs across AI Overviews rollouts | About 15% less daily traffic to exposed English articles |
| Wikimedia Foundation, October 2025 | Human pageviews March to August 2025 vs 2024 | Down about 8% after bot reclassification |
| Cloudflare Radar, June 2026 | Share of HTML requests from bots vs humans | 57.3% bots, 42.7% humans |
| Steinacker-Olsztyn, Gosain and Dao, WWW 2026 | robots.txt rules on reputable vs misinformation sites | 60.0% vs 9.1% disallow at least one AI crawler |
| Grossman et al., SIGIR 2026 | 11,500 queries across Search, AI Overviews and Gemini | Sites blocking the AI crawler are significantly less likely to be cited |
| Allaham and Diakopoulos, 2026 | 712 queries across four generative search engines | About 16% of cited sources are AI-generated sites |
| TollBit via Digiday, 2026 | AI bot scrapes against robots.txt permissions, Q4 2025 | 30% did not abide by the rules |
What Pay-to-Crawl and Licensing Deals Would Change
If blocking is the stick, payment is the carrot that has so far failed to appear. The RMIT authors note that “pay to crawl” models “have been suggested as a way to compensate content creators, but haven’t gained traction”. The infrastructure exists; the buyers mostly do not.
From Pay Per Crawl to Pay Per Use
Cloudflare launched its Pay Per Crawl marketplace in July 2025 alongside the one-click “Block AI Bots” option. On 1 July 2026 it said the model was evolving into “Pay Per Use”, where a publisher is paid when its content creates value in an AI product, not merely when it is fetched. The first partners are Ceramic.ai and You.com: an opted-in publisher is paid when its content appears in Ceramic’s AI search results or when You.com accesses a premium article. Other AI companies “can customize this model”, Cloudflare says, which is a polite way of noting that none of the largest ones has signed up.
Blocking as a negotiating tactic
The publishers interviewed by Digiday were candid that blocking AI crawlers is mainly leverage. Alphonse Hardel of Reuters called robots.txt “non-binding” but said it “strengthens a message to say, if you want this, let’s have a conversation”. Lindsay Van Kirk of People Inc. argued that “if we don’t put friction into this side of the economy now, it makes it harder for us to recapture value on the other side of it”. Reuters already holds AI licensing agreements, and it credits friction with bringing AI companies to the table.
Why smaller sites are stuck
Licensing works for Reuters. It does not work for a regional news site, a specialist blog or a company knowledge base, none of which can negotiate with an AI lab. Cloudflare’s own blog describes the “Faustian bargain” facing a small site: “either show up in search and let AI train on you, or risk losing discoverability”. That is why the network-level default matters more than any individual deal, and why the sites that produce the reliable niche information AI answers most need are the ones most likely to vanish from them.
What It Means for You: Finding Reliable Information Around AI Crawlers
The RMIT authors’ advice to readers is plain: “The quality of AI summaries is likely to go down, at least in the short term, while the new economics of the web get sorted out.” Two forces drive that. High-quality content is less likely to feed the summaries, and models trained on more AI-generated text may degrade. Their practical suggestions follow from that diagnosis.
Scroll down and click
“For now, whatever search engine you’re using, the best thing you can do is to scroll down and click on some actual search results,” they write. “This benefits content creators, and is also more likely to give you more accurate information.” The Pew data shows how rare that has become: one click in twelve visits when a summary is present. Every click is a vote for the site that did the work, and a signal to the platform that referrals still matter.
Search engines that do not summarise
The authors point to search engines without AI summaries as a way to sidestep the problem. ZDNet recommends Mojeek, PCMag recommends Brave, and Ban the Bots lists several, including small-web engines such as Marginalia and Wiby aimed at “small producer” content like blogs. Ban the Bots is careful to note that “without AI” usually means link-first results with no AI summary, not a search engine that uses no machine learning at all.
Check what the AI actually cited
When a chatbot does cite sources, open them. The Allaham and Diakopoulos audit found generative engines repeatedly cite a narrow set of domains while surfacing a long tail of minimally cited ones, and roughly one in six of those sources is AI-generated. This site’s guide to detecting AI writing covers the tells and, more importantly, their limits.
| Option | Who recommends it | What you get | Trade-off |
|---|---|---|---|
| Mojeek | ZDNet; RMIT authors | Independent index, no AI summary, UK-based | Smaller index than Google |
| Brave Search | PCMag; RMIT authors | Independent index, AI answers can be switched off | AI features are on by default |
| Marginalia and Wiby | Ban the Bots | Small-web discovery, blogs and personal sites | Not built for everyday queries |
| Google with AI Overviews reduced | Ban the Bots | Familiar results with fewer summaries | Cannot be fully disabled in 2026 |
| Scroll past the summary and click | RMIT authors | Original source, full context, supports the publisher | Takes longer than reading the summary |
What UK Businesses Should Do About AI Crawlers on Their Own Website
Most companies are on both sides of this story. They want customers to find them, which increasingly means being cited by an AI answer, and they do not want their product documentation, pricing pages and expertise absorbed for nothing. The decisions below are the ones worth making before Cloudflare’s defaults change on 15 September.
Decide per bot class, not for “AI” as a whole
Cloudflare’s Search, Agent and Training split is a useful frame even if you are not on Cloudflare. Search crawlers are the ones that can send a citation and a visitor; an Agent bot may be a prospect’s assistant checking your opening hours; a Training crawler gives nothing back. Blocking all three because the first one annoys you is how a company disappears from the answers its buyers now read. A technical SEO audit should now include a bot-traffic review and a deliberate robots.txt policy for AI crawlers.
Make the content worth citing and measure whether it is
The SIGIR study shows generative search favours a different set of sources from classic search, so ranking well on Google no longer guarantees a citation. This site’s guides to citable content for AI search and to measuring brand visibility in ChatGPT, Gemini and Perplexity cover both halves, and the GEO reporting framework turns mentions and citations into numbers a board will accept. AEO services exist precisely because the answer box is now the first result.
Watch the liability side
Being cited is not always good news. As this site reported when a court ruled Google liable for false statements in AI Overviews, summaries can misstate what a company said. If AI crawlers reach a stale pricing page or an outdated policy, the summary will repeat it. Keep the pages AI crawlers can reach current, and use structured data so the machine-readable version matches the human one. The AI models and tools hub tracks how each engine sources its answers as the picture changes.
Do not build your own traffic forecast on last year’s search referrals
The Wikipedia data is the cleanest evidence that referral traffic falls the moment AI summaries arrive for a query type, and the drop is largest for informational content. Marketing plans that assume steady organic traffic from explainer content need revising, which is a marketing services conversation as much as a technical one. Independent indexes built for AI agents, such as the web index Keenable is building, are another channel to watch, because they decide which sites an agent sees at all.
Frequently Asked Questions
Who wrote “AI is eating website traffic, websites are blocking AI”?
Dana McKay and Damiano Spina of RMIT University in Melbourne. It was published by The Conversation on 31 August 2026 under a Creative Commons licence and syndicated the same day by TechXplore and several other outlets.
Are AI crawlers really more than half of web traffic?
Cloudflare reported on 5 June 2026 that bots of all kinds made up 57.3 percent of HTML requests on its network. That includes search crawlers and other automation as well as AI crawlers, so “over half is AI bots” is a rounding of “over half is automated”, with AI-driven agents the fastest-growing part.
What happens on 15 September 2026?
Cloudflare’s new defaults take effect. On pages that display ads, Training and Agent crawlers are blocked by default and Search crawlers remain allowed, for new domains and, per TechCrunch, existing free-tier customers. Multi-purpose crawlers such as Googlebot are judged by their most restrictive behaviour. Owners can opt out.
Does blocking AI crawlers stop AI from using my content?
Not entirely. robots.txt is voluntary, TollBit found 30 percent of scrapes ignored it in late 2025, and the SIGIR study shows AI Overviews can still ground answers on content it has access to. Blocking mainly reduces how often you are cited and creates leverage for a licensing conversation.
Why does blocking make AI answers less reliable?
Because blocking is lopsided. Sixty percent of reputable news sites disallow at least one AI crawler while 9.1 percent of misinformation sites do, so the sources still open to AI crawlers skew towards the unreliable, and about one in six sources cited by AI search is already AI-generated.
What is the simplest thing an individual can do?
Scroll past the AI summary and click a real result, as the RMIT authors advise, or use a search engine without AI summaries such as Mojeek or Brave for research that matters.
References
Impact of AI Search Summaries on Website Traffic: Evidence from Google AI Overviews and Wikipedia
New User Trends on Wikipedia (Wikimedia Foundation)
Google Zero is here, and it is a disaster for the open web (The Verge)
Cloudflare: Bots now make up 57% of webpage requests (Search Engine Land)
Your site, your rules: new AI traffic options for all customers (Cloudflare Blog)
Cloudflare’s new policy pushes AI companies to pay for publishers’ content (TechCrunch)
Reuters and Time adopt bot-blocking whitelists to rein in AI crawlers (Digiday)
Multiple AI companies bypassing web standard to scrape publisher sites (Reuters)
Is Misinformation More Open? A Study of robots.txt Gatekeeping on the Web
Synthetic Sources? Auditing Generative Search Engine Citations for Evidence of AI-Generated Sources
AI models collapse when trained on recursively generated data (Nature)
Google search alternatives without AI (ZDNet)
Search Engines Without AI: 8 Privacy-First Picks (Ban the Bots)
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.