searchDecoded Research · UK consumer AI controls study 2026
UK consumer AI crawler and content-use study 2026
Only 13% of prominent UK consumer websites studied explicitly address AI-specific controls
A robots.txt file publishes instructions for automated crawlers. searchDecoded selected the 100 websites and defined the eight AI crawler and content-use controls before collecting any policy data.
Key findings
Explicit AI-specific rules were uncommon
13 of 100 sites explicitly named at least one of the eight AI crawler or content-use controls in the main analysis.
72 of 100 sites applied their general robots rules to AI-related controls rather than publishing an AI-specific rule.
8 of 100 sites explicitly named at least one search control, compared with 12 of 100 sites naming at least one training or content-use control.
Finding 1
General robots rules applied more often than AI-specific rules
Only 13% of sites explicitly named an AI-specific control. On 72%, AI-related controls instead inherited the website's general robots rules.
Explicit AI-specific rules name at least one of the eight controls in the main analysis.
Inherited general rules means the site didn't name an AI-specific control, so its general robots rules applied.
Unavailable or unclassified means the collected evidence couldn't support a reliable result.
No matching robots group means the published file contained no group that applied to the control.
No robots policy means no usable file was published at the standard location.
How sites published AI-related robots policy
The chart shows the different robots policy states found across the 100 sites.
Source: searchDecoded, UK consumer AI crawler and content-use study 2026.
View data table
Scroll sideways to see every column.
| Category | Result | Percentage |
|---|---|---|
| Explicit AI-specific rules | 13/100 | 13% |
| AI controls inherited general rules | 72/100 | 72% |
| Unavailable or could not be classified | 5/100 | 5% |
| No matching robots group | 7/100 | 7% |
| No robots policy | 3/100 | 3% |
Finding 2
Training controls were more often named on their own
OpenAI and Anthropic publish separate controls for search and model development. These charts show which controls sites named, not whether those controls were allowed or blocked.
OpenAI naming pair
GPTBot is the model-development control; OAI-SearchBot supports ChatGPT search and discovery.
Source: searchDecoded, UK consumer AI crawler and content-use study 2026.
View data table
Scroll sideways to see every column.
| Category | Result | Percentage |
|---|---|---|
| training only | 7/100 | 7% |
| search only | 1/100 | 1% |
| both | 5/100 | 5% |
| neither | 87/100 | 87% |
Anthropic naming pair
ClaudeBot is the model-development crawler; Claude-SearchBot supports search and retrieval.
Source: searchDecoded, UK consumer AI crawler and content-use study 2026.
View data table
Scroll sideways to see every column.
| Category | Result | Percentage |
|---|---|---|
| training only | 8/100 | 8% |
| search only | 1/100 | 1% |
| both | 4/100 | 4% |
| neither | 87/100 | 87% |
Finding 3
Sites named AI-specific controls in groups
Among sites that named AI-specific controls, none named just one, two or three.
How many controls each site named
Source: searchDecoded, UK consumer AI crawler and content-use study 2026.
View data table
Scroll sideways to see every column.
| Category | Sites | Percentage |
|---|---|---|
| 0 controls | 87/100 | 87% |
| 1 control | 0/100 | 0% |
| 2 controls | 0/100 | 0% |
| 3 controls | 0/100 | 0% |
| 4 controls | 8/100 | 8% |
| 5 controls | 0/100 | 0% |
| 6 controls | 1/100 | 1% |
| 7 controls | 2/100 | 2% |
| 8 controls | 2/100 | 2% |
Exploratory findingApplebot-Extended, ClaudeBot, GPTBot and Google-Extended appeared together on five of the 13 sites that explicitly named AI controls (38.5%).
Finding 4
AI-specific controls by sector
Sector sample sizes vary and are relatively small, and the study isn't weighted to represent the wider UK consumer economy.
Control detail
Which AI-specific controls sites named
The eight controls aren't all used for the same purpose. Some relate to AI search or model development, while others govern how previously collected content can be used.
Scroll sideways to see every column.
| Control | Technical role | Explicitly named | Indeterminate |
|---|---|---|---|
| GPTBot | Crawler Model development/training | 12% | 5% |
| Google-Extended | Content-use token Model development/grounding | 12% | 5% |
| ClaudeBot | Crawler Model development/training | 12% | 5% |
| PerplexityBot | Crawler AI search/retrieval | 8% | 5% |
| Applebot-Extended | Content-use token Model development/training | 8% | 5% |
| OAI-SearchBot | Crawler AI search/retrieval | 6% | 5% |
| Claude-SearchBot | Crawler AI search/retrieval | 5% | 5% |
| Claude-User | User-triggered retrieval User-triggered retrieval | 5% | 5% |
What the explicitly named rules actually did
Explicitly addressing an AI-specific control didn't automatically mean blocking it. The published rules included full-site allows, path-specific restrictions and sitewide blocks. This distinction matters because explicitly naming a control doesn't reveal the direction of the policy on its own. Percentages below use all 100 sites.
- Explicit allow
- The named control's applicable rules permitted access or use across the site.
- Partial restriction
- Some paths were restricted, but the entire site was not blocked.
- Sitewide block
- The applicable policy disallowed the named control from
/across the site.
Scroll sideways to see every column.
| Control | Explicitly named | Explicit allow | Partial restriction | Sitewide block |
|---|---|---|---|---|
| GPTBot | 12% · 12 sites | 3% · 3 sites | 4% · 4 sites | 5% · 5 sites |
| Google-Extended | 12% · 12 sites | 3% · 3 sites | 4% · 4 sites | 5% · 5 sites |
| ClaudeBot | 12% · 12 sites | 3% · 3 sites | 4% · 4 sites | 5% · 5 sites |
| PerplexityBot | 8% · 8 sites | 3% · 3 sites | 4% · 4 sites | 1% · 1 site |
| Applebot-Extended | 8% · 8 sites | 1% · 1 site | 3% · 3 sites | 4% · 4 sites |
| OAI-SearchBot | 6% · 6 sites | 1% · 1 site | 4% · 4 sites | 1% · 1 site |
| Claude-SearchBot | 5% · 5 sites | 0% · 0 sites | 4% · 4 sites | 1% · 1 site |
| Claude-User | 5% · 5 sites | 0% · 0 sites | 4% · 4 sites | 1% · 1 site |
How robots directives workAllow: / can produce an explicit full-site allow, while Disallow: / can produce a sitewide block. Combinations of path-specific rules can result in a partial restriction. Each robots.txt file was assessed as a whole using the Robots Exclusion Protocol (RFC 9309), rather than judging individual directives in isolation.
Additional context
Six additional controls provided context
Googlebot, Applebot and bingbot showed how sites treated ordinary search crawlers. ChatGPT-User and Perplexity-User fetch pages in response to a user's request, while CCBot collects web data for Common Crawl. These six controls were analysed separately from the eight AI-specific controls in the main findings.
How this compares with other researchA 2025 Cloudflare study found AI-specific directives on around 14% of a different group of top websites. Its sample, crawler definitions and methodology differed from this study, so the figures aren't directly comparable, but it provides useful wider context.
Reading the findings
What these results do not mean
- They don't test whether crawlers follow robots.txt instructions.
- They don't show whether a site appears in, or is cited by, an AI answer.
- They don't establish why an organisation published a rule, or whether that rule is good or bad.
- They don't grade a business as ‘AI ready’, ‘AI friendly’ or ‘AI hostile’.
- They don't prove that applying general rules to AI controls was deliberate, or that a restriction prevents every possible form of AI access.
Methodology
How the study was conducted
searchDecoded selected the 100 websites and defined the crawler and content-use controls before collecting any policy data. This helped prevent the results from influencing the method.
Sampling
searchDecoded deliberately selected 100 prominent consumer-facing websites across six sectors. This is called a purposive stratified sample. Independent regulatory, consumer-watchdog and industry data defined which sites belonged in each sector, and searchDecoded locked the list before collecting robots policy.
Controls and collection
searchDecoded locked the eight AI crawler and content-use controls before collection. On 18 August 2026, searchDecoded collected each available website's root robots.txt file once, then checked all eight controls against the preserved evidence.
Classification
The same automated rules classified every website under RFC 9309 and the provider-specific rules defined in advance. searchDecoded kept explicit rules, inherited rules, missing files, retrieval failures and unclear results separate instead of treating ‘not blocked’ as ‘allowed’.
Retrieval quality assurance
When automated retrieval failed, a predefined manual browser pass recovered usable root robots evidence for 10 sites through normal browser navigation. Five sites remained unresolved. Alternative hostnames and paths weren't substituted, the recovered files used the unchanged classifier, and no missing values were estimated.
Limitations
- The deliberately selected sample isn't representative of all UK businesses or websites.
- The six sector groups have different sizes, so sector findings are descriptive.
- One consumer domain represents each sampled organisation or brand.
- This is a point-in-time snapshot; policies and provider documentation can change.
- Published instructions show policy, not whether a crawler follows it or why an organisation chose it.
- Google-Extended and Applebot-Extended are content-use controls, not HTTP crawlers.
- searchDecoded didn't estimate results for retrieval failures or observations that remained indeterminate after review.
- Partial restrictions may apply only to particular paths.
Open data
Download the study data
These versioned files were generated from the locked observations and analysis, so their figures match the report. The raw robots.txt evidence remains private.
Data license
The original and derived datasets published with this study are licensed under Creative Commons Attribution 4.0 International (CC BY 4.0). Please credit searchDecoded and link to this study when reusing the data.
The licence applies only to searchDecoded's original and derived datasets published with this study. It does not apply to third-party robots.txt content, source material, company names, logos or trademarks. searchDecoded branding, website copy, design and other original creative assets remain subject to their existing rights and are not licensed for reuse under CC BY 4.0.
Suggested attribution
Source: searchDecoded, ‘UK consumer AI crawler and content-use study 2026’.
This is a suggested format, not the only valid form of attribution.
Analysis version: v2. Technical checksums are included with the downloadable data.
Citation
Cite this study
Use the formats below when referencing this study in an article, report or dataset. Each citation identifies the published version and canonical report URL.
Choose a citation formatCopy the study reference as Web, Markdown, HTML or Harvard.
Citation conventions vary between publishers and institutions. Check any house style before submitting.