searchDecoded Research · UK consumer AI controls study 2026

UK consumer AI crawler and content-use study 2026

Only 13% of prominent UK consumer websites studied explicitly address AI-specific controls

by searchDecoded100 sites analysed · six UK consumer sectors

A robots.txt file publishes instructions for automated crawlers. searchDecoded selected the 100 websites and defined the eight AI crawler and content-use controls before collecting any policy data.

Key findings

Explicit AI-specific rules were uncommon

13%

13 of 100 sites explicitly named at least one of the eight AI crawler or content-use controls in the main analysis.

72%

72 of 100 sites applied their general robots rules to AI-related controls rather than publishing an AI-specific rule.

8% vs 12%

8 of 100 sites explicitly named at least one search control, compared with 12 of 100 sites naming at least one training or content-use control.

Finding 1

General robots rules applied more often than AI-specific rules

Only 13% of sites explicitly named an AI-specific control. On 72%, AI-related controls instead inherited the website's general robots rules.

Explicit AI-specific rules name at least one of the eight controls in the main analysis.

Inherited general rules means the site didn't name an AI-specific control, so its general robots rules applied.

Unavailable or unclassified means the collected evidence couldn't support a reliable result.

No matching robots group means the published file contained no group that applied to the control.

No robots policy means no usable file was published at the standard location.

How sites published AI-related robots policy

The chart shows the different robots policy states found across the 100 sites.

13%Explicit AI-specific rules13 sites
72%AI controls inherited general rules72 sites
5%Unavailable or could not be classified5 sites
7%No matching robots group7 sites
3%No robots policy3 sites

Source: searchDecoded, UK consumer AI crawler and content-use study 2026.

View data table

Scroll sideways to see every column.

CategoryResultPercentage
Explicit AI-specific rules13/10013%
AI controls inherited general rules72/10072%
Unavailable or could not be classified5/1005%
No matching robots group7/1007%
No robots policy3/1003%

Finding 2

Training controls were more often named on their own

OpenAI and Anthropic publish separate controls for search and model development. These charts show which controls sites named, not whether those controls were allowed or blocked.

OpenAI naming pair

GPTBot is the model-development control; OAI-SearchBot supports ChatGPT search and discovery.

training only: 7 of 100 (7%)search only: 1 of 100 (1%)both: 5 of 100 (5%)neither: 87 of 100 (87%)
7%training only7 sites
1%search only1 site
5%both5 sites
87%neither87 sites

Source: searchDecoded, UK consumer AI crawler and content-use study 2026.

View data table

Scroll sideways to see every column.

CategoryResultPercentage
training only7/1007%
search only1/1001%
both5/1005%
neither87/10087%

Anthropic naming pair

ClaudeBot is the model-development crawler; Claude-SearchBot supports search and retrieval.

training only: 8 of 100 (8%)search only: 1 of 100 (1%)both: 4 of 100 (4%)neither: 87 of 100 (87%)
8%training only8 sites
1%search only1 site
4%both4 sites
87%neither87 sites

Source: searchDecoded, UK consumer AI crawler and content-use study 2026.

View data table

Scroll sideways to see every column.

CategoryResultPercentage
training only8/1008%
search only1/1001%
both4/1004%
neither87/10087%

Finding 3

Sites named AI-specific controls in groups

Among sites that named AI-specific controls, none named just one, two or three.

How many controls each site named

0 controls
87%87 sites
1 control
0%0 sites
2 controls
0%0 sites
3 controls
0%0 sites
4 controls
8%8 sites
5 controls
0%0 sites
6 controls
1%1 site
7 controls
2%2 sites
8 controls
2%2 sites

Source: searchDecoded, UK consumer AI crawler and content-use study 2026.

View data table

Scroll sideways to see every column.

CategorySitesPercentage
0 controls87/10087%
1 control0/1000%
2 controls0/1000%
3 controls0/1000%
4 controls8/1008%
5 controls0/1000%
6 controls1/1001%
7 controls2/1002%
8 controls2/1002%

Exploratory findingApplebot-Extended, ClaudeBot, GPTBot and Google-Extended appeared together on five of the 13 sites that explicitly named AI controls (38.5%).

Finding 4

AI-specific controls by sector

Sector sample sizes vary and are relatively small, and the study isn't weighted to represent the wider UK consumer economy.

Explicit AI-specific naming by sector

Grocery retail
21.4%3 of 14
Package travel
20%3 of 15
Personal banking
11.8%2 of 17
Domestic energy
12.5%2 of 16
Passenger rail
5%1 of 20
Automotive
11.1%2 of 18

Source: searchDecoded, UK consumer AI crawler and content-use study 2026.

View data table

Scroll sideways to see every column.

CategoryResultPercentage
Grocery retail3/1421.4%
Package travel3/1520%
Personal banking2/1711.8%
Domestic energy2/1612.5%
Passenger rail1/205%
Automotive2/1811.1%

Control detail

Which AI-specific controls sites named

The eight controls aren't all used for the same purpose. Some relate to AI search or model development, while others govern how previously collected content can be used.

How often each control was named

GPTBotCrawler
12%12 sites
Google-ExtendedContent-use token
12%12 sites
ClaudeBotCrawler
12%12 sites
PerplexityBotCrawler
8%8 sites
Applebot-ExtendedContent-use token
8%8 sites
OAI-SearchBotCrawler
6%6 sites
Claude-SearchBotCrawler
5%5 sites
Claude-UserUser-triggered retrieval
5%5 sites

Source: searchDecoded, UK consumer AI crawler and content-use study 2026.

View data table

Scroll sideways to see every column.

CategoryResultPercentageTechnical role
GPTBot12/10012%Crawler
Google-Extended12/10012%Content-use token
ClaudeBot12/10012%Crawler
PerplexityBot8/1008%Crawler
Applebot-Extended8/1008%Content-use token
OAI-SearchBot6/1006%Crawler
Claude-SearchBot5/1005%Crawler
Claude-User5/1005%User-triggered retrieval

Scroll sideways to see every column.

ControlTechnical roleExplicitly namedIndeterminate
GPTBotCrawler
Model development/training
12%5%
Google-ExtendedContent-use token
Model development/grounding
12%5%
ClaudeBotCrawler
Model development/training
12%5%
PerplexityBotCrawler
AI search/retrieval
8%5%
Applebot-ExtendedContent-use token
Model development/training
8%5%
OAI-SearchBotCrawler
AI search/retrieval
6%5%
Claude-SearchBotCrawler
AI search/retrieval
5%5%
Claude-UserUser-triggered retrieval
User-triggered retrieval
5%5%

What the explicitly named rules actually did

Explicitly addressing an AI-specific control didn't automatically mean blocking it. The published rules included full-site allows, path-specific restrictions and sitewide blocks. This distinction matters because explicitly naming a control doesn't reveal the direction of the policy on its own. Percentages below use all 100 sites.

Explicit allow
The named control's applicable rules permitted access or use across the site.
Partial restriction
Some paths were restricted, but the entire site was not blocked.
Sitewide block
The applicable policy disallowed the named control from / across the site.

Scroll sideways to see every column.

ControlExplicitly namedExplicit allowPartial restrictionSitewide block
GPTBot12% · 12 sites3% · 3 sites4% · 4 sites5% · 5 sites
Google-Extended12% · 12 sites3% · 3 sites4% · 4 sites5% · 5 sites
ClaudeBot12% · 12 sites3% · 3 sites4% · 4 sites5% · 5 sites
PerplexityBot8% · 8 sites3% · 3 sites4% · 4 sites1% · 1 site
Applebot-Extended8% · 8 sites1% · 1 site3% · 3 sites4% · 4 sites
OAI-SearchBot6% · 6 sites1% · 1 site4% · 4 sites1% · 1 site
Claude-SearchBot5% · 5 sites0% · 0 sites4% · 4 sites1% · 1 site
Claude-User5% · 5 sites0% · 0 sites4% · 4 sites1% · 1 site

How robots directives workAllow: / can produce an explicit full-site allow, while Disallow: / can produce a sitewide block. Combinations of path-specific rules can result in a partial restriction. Each robots.txt file was assessed as a whole using the Robots Exclusion Protocol (RFC 9309), rather than judging individual directives in isolation.

Additional context

Six additional controls provided context

Googlebot, Applebot and bingbot showed how sites treated ordinary search crawlers. ChatGPT-User and Perplexity-User fetch pages in response to a user's request, while CCBot collects web data for Common Crawl. These six controls were analysed separately from the eight AI-specific controls in the main findings.

How this compares with other researchA 2025 Cloudflare study found AI-specific directives on around 14% of a different group of top websites. Its sample, crawler definitions and methodology differed from this study, so the figures aren't directly comparable, but it provides useful wider context.

Reading the findings

What these results do not mean

  • They don't test whether crawlers follow robots.txt instructions.
  • They don't show whether a site appears in, or is cited by, an AI answer.
  • They don't establish why an organisation published a rule, or whether that rule is good or bad.
  • They don't grade a business as ‘AI ready’, ‘AI friendly’ or ‘AI hostile’.
  • They don't prove that applying general rules to AI controls was deliberate, or that a restriction prevents every possible form of AI access.

Methodology

How the study was conducted

searchDecoded selected the 100 websites and defined the crawler and content-use controls before collecting any policy data. This helped prevent the results from influencing the method.

Sampling

searchDecoded deliberately selected 100 prominent consumer-facing websites across six sectors. This is called a purposive stratified sample. Independent regulatory, consumer-watchdog and industry data defined which sites belonged in each sector, and searchDecoded locked the list before collecting robots policy.

Controls and collection

searchDecoded locked the eight AI crawler and content-use controls before collection. On 18 August 2026, searchDecoded collected each available website's root robots.txt file once, then checked all eight controls against the preserved evidence.

Classification

The same automated rules classified every website under RFC 9309 and the provider-specific rules defined in advance. searchDecoded kept explicit rules, inherited rules, missing files, retrieval failures and unclear results separate instead of treating ‘not blocked’ as ‘allowed’.

Retrieval quality assurance

When automated retrieval failed, a predefined manual browser pass recovered usable root robots evidence for 10 sites through normal browser navigation. Five sites remained unresolved. Alternative hostnames and paths weren't substituted, the recovered files used the unchanged classifier, and no missing values were estimated.

Limitations

  • The deliberately selected sample isn't representative of all UK businesses or websites.
  • The six sector groups have different sizes, so sector findings are descriptive.
  • One consumer domain represents each sampled organisation or brand.
  • This is a point-in-time snapshot; policies and provider documentation can change.
  • Published instructions show policy, not whether a crawler follows it or why an organisation chose it.
  • Google-Extended and Applebot-Extended are content-use controls, not HTTP crawlers.
  • searchDecoded didn't estimate results for retrieval failures or observations that remained indeterminate after review.
  • Partial restrictions may apply only to particular paths.

Open data

Download the study data

These versioned files were generated from the locked observations and analysis, so their figures match the report. The raw robots.txt evidence remains private.

Data license

The original and derived datasets published with this study are licensed under Creative Commons Attribution 4.0 International (CC BY 4.0). Please credit searchDecoded and link to this study when reusing the data.

The licence applies only to searchDecoded's original and derived datasets published with this study. It does not apply to third-party robots.txt content, source material, company names, logos or trademarks. searchDecoded branding, website copy, design and other original creative assets remain subject to their existing rights and are not licensed for reuse under CC BY 4.0.

Suggested attribution

Source: searchDecoded, ‘UK consumer AI crawler and content-use study 2026’.

This is a suggested format, not the only valid form of attribution.

Analysis version: v2. Technical checksums are included with the downloadable data.

Citation

Cite this study

Use the formats below when referencing this study in an article, report or dataset. Each citation identifies the published version and canonical report URL.

Choose a citation formatCopy the study reference as Web, Markdown, HTML or Harvard.
Citation format

Citation conventions vary between publishers and institutions. Check any house style before submitting.