The AI Hypocrisy Index measures the gap between what a company writes in its Terms of Service and what it does with data. We scored 35 AI companies and the 150 most-visited consumer websites on two 0-to-10 scales. The forbid score measures how hard the Terms ban scraping. The take score measures how broadly the company claims user or public data, and their product divided by ten is the Index. It only climbs when a company scores high on both.
Adobe, Perplexity, and LinkedIn top the Index. All three ban scraping in strong Terms language but leave every AI crawler unblocked in their robots.txt. Five of the top ten companies follow that pattern.
19 of 35 AI companies write an explicit anti-scraping ban into their Terms while claiming rights to user or public data. 85 of the 150 top consumer sites do the same. 47 times a platform’s robots.txt disallows a rival AI company’s crawler.
X Corp charges $15,000 (and €15,000) per million posts scraped. Nine companies train on user content by default, requiring an explicit opt-out. Adobe scores 4.4 out of 10 and holds that position under every alternative weighting scheme we tried. 17 firms share nearly identical “robot, spider, scraper” phrasing across their Terms of Service.
This is the July 2026 snapshot. Every claim is backed by an exact-substring-verified quote from the source document.
What we measured
We split the sample into two tiers. Tier A is 35 companies that build AI models or operate crawlers. For each we read the Terms of Service, Privacy Policy, and robots.txt. Tier B is the 150 most-visited consumer websites drawn from the Tranco top ranking (list 94Y82). For Tier B we read Terms of Service and robots.txt only. HasData and its own domain were excluded from the sample.
Documents were fetched on 23 July 2026 with a custom User-Agent (DoubleStandardIndexResearch/1.0), at one request per second per domain. No paywalls, logins, or bot protections were circumvented.
Each document was classified against a 12-clause taxonomy, split into 7 anti-scraping clauses and 5 extraction clauses. An LLM classified each document at temperature zero, returning strict JSON. Every reported clause carries a verbatim quotation programmatically verified as an exact substring of the source document. Classification was run twice. 402 findings passed at a 100% exact-substring rate. 12 findings failed verification and were dropped.
Each Tier-A company gets two scores on a 0 to 10 scale. The forbid score counts anti-scraping clauses. The take score counts extraction clauses. The AI Hypocrisy Index is their product divided by ten. A pure blocker or a pure scraper stays low. The Index climbs only when a company scores high on both sides.
The numbers at a glance:
| Metric | Tier A (35 companies) | Tier B (150 top sites) |
|---|---|---|
| Bans scraping in Terms | 19 | 85 |
| Requires prior written consent | 6 | 11 |
| Explicit AI-training ban on their own content | 5 | 22 |
| Blocks 1+ AI training crawlers in robots.txt | 8 | 44 |
Privacy Policies were not read for Tier B, so the following apply to Tier A only. 9 companies train on user content by default (opt-out). 7 companies claim rights to “publicly available” web data. 1 company has a liquidated-damages clause (X Corp).
The measurements use these definitions.
- Anti-scraping ban. Terms of Service contain at least one of the 7 AS clauses (general scraping ban, absolute ban, prior written consent, AI-training ban on their content, liquidated damages, anti-circumvention, API-only access).
- AI training crawler. A user-agent operated by an AI vendor for training-data collection. 13 were tracked, including
GPTBot,ClaudeBot,Google-Extended,Applebot-Extended,PerplexityBot,CCBot, andMeta-ExternalAgent. - Paper-only forbidder. Terms declare an anti-scraping stance (forbid score ≥ 2.5) and robots.txt blocks zero of the 13 tracked AI training crawlers.
- Silent enforcer. Terms are permissive (forbid score < 2.5) and robots.txt blocks at least one AI training crawler.
- Verbatim verification. A reported clause is included only if its exact wording is found as an exact substring of the fetched source document.
One of 35 Tier-A Terms pages was a JavaScript-only site and could not be read in text form. The historical (Wayback) dimension is partial and is used for context only.
The AI Hypocrisy Index leaderboard
Adobe, Perplexity, and LinkedIn top the list. Every one of them tells you not to scrape their content and blocks zero AI training crawlers in robots.txt.
Adobe processes your documents and images through AI, claims rights over what you upload, and tells everyone else that scraping Adobe is not allowed. Forbid 5.5, take 8.
Perplexity scored 6.5 on the forbid component, the highest of any company in the sample that built its product by indexing the web. Its robots.txt blocks zero AI training crawlers.

Table 1 lists the top 15 Tier-A companies with their forbid score, take score, composite Index, and the number of AI training crawlers each blocks in its own robots.txt.
| # | Company | Domain | Forbid | Take | Index | AI bots blocked |
|---|---|---|---|---|---|---|
| 1 | Adobe | adobe.com | 5.5 | 8 | 4.4 | 0 |
| 2 | Perplexity | perplexity.ai | 6.5 | 6 | 3.9 | 0 |
| 3 | linkedin.com | 6.5 | 4 | 2.6 | 0 | |
| 4 | reddit.com | 4 | 6 | 2.4 | 11 | |
| 5 | Yelp | yelp.com | 3 | 8 | 2.4 | 11 |
| 6 | Amazon | amazon.com | 3.5 | 6 | 2.1 | 9 |
| 7 | X Corp | x.com | 5.0 | 4 | 2.0 | 10 |
| 8 | OpenAI | openai.com | 2.5 | 7 | 1.75 | 0 |
| 9 | google.com | 2.5 | 7 | 1.75 | 0 | |
| 10 | Meta | facebook.com | 2.5 | 6 | 1.5 | 11 |
| 11 | Cohere | cohere.com | 2.5 | 6 | 1.5 | 0 |
| 12 | Snap | snap.com | 2.5 | 6 | 1.5 | 0 |
| 13 | Midjourney | midjourney.com | 4.5 | 3 | 1.35 | 0 |
| 14 | Expedia | expedia.com | 4.5 | 3 | 1.35 | 0 |
| 15 | Airbnb | airbnb.com | 4.5 | 3 | 1.35 | 0 |
Anthropic trains on your Claude conversations by default. Take 5, forbid 0, composite 0.0. GitHub claims rights over code in public repositories. Take 9, the highest in the sample, forbid 0, composite 0.0.
The forbid-vs-take split
The content platforms own the upper-right of this chart. Adobe, Perplexity, Reddit, and Yelp all restrict access aggressively in their Terms and claim broad rights over the content on their platforms.

OpenAI and Google both score 2.5 on forbid and 7 on take, claim rights to “publicly available” data, and run training crawlers.
X Corp and LinkedIn score high on forbid (5.0 and 6.5) but only 4 on take. They enforce access restrictions but claim less from user content than Reddit or Yelp.
Anthropic, Microsoft, Apple, Nvidia, Salesforce/Slack, xAI, Mistral, Stability AI, and Stack Overflow all score 0 on forbid. They collect data and place no legal restrictions on scraping. Composite 0.0.
Banning scraping is the norm
19 of the 35 AI companies write an explicit anti-scraping clause into their Terms of Service. Across the 150 most-visited consumer websites, 85 do the same. It’s the default across the top sites in the sample.

The AI builders are in this group too. Google’s Terms prohibit “scraping content that doesn’t belong to you.” OpenAI bans “automatically or programmatically extracting data or Output.” Both companies also describe collecting “information that is publicly available on the internet” to train their models.
Seven companies banned scraping in their Terms and claimed rights to public web data in their Privacy Policy. Adobe, Booking, Getty Images, Google, LinkedIn, OpenAI, and Yelp. Both documents are on the same domain.
AI labs block each other in robots.txt
Reading each AI platform’s robots.txt against every other AI company’s crawler user-agent, we found 47 instances of a platform disallowing a rival AI training crawler. Reddit and Yelp each block 11 of the 13 tracked AI training bots. Meta blocks 11. X Corp, Pinterest, and Quora each block 10. Amazon blocks 9. ByteDance/TikTok blocks 8.

Which platform blocks which rival is visible in the matrix. X Corp’s robots.txt disallows OpenAI, Anthropic, Google, Perplexity, and six others. Meta disallows ten rivals. The refusal is mutual across almost every pair of platforms in the sample.

Every AI vendor we tracked appears on someone’s disallow list. OpenAI, Anthropic, Google, Perplexity, Apple, Amazon, Meta, ByteDance, Mistral, Cohere, Common Crawl, and Microsoft all get blocked by at least one platform. The phrase “publicly available” describes data that is available to everyone except the specific crawlers each platform wants to keep out.
Where robots.txt disagrees with Terms
The Terms of Service and the robots.txt are separate signals from the same company. In this sample, they disagree in both directions.
| robots.txt blocks 1+ AI training crawler | robots.txt blocks 0 AI training crawlers | |
|---|---|---|
| Terms forbid score ≥ 2.5 | Aligned strict. X Corp, Reddit, Amazon, Yelp, Meta, TikTok. | Paper-only forbidders. Adobe, Perplexity, LinkedIn, Midjourney, Expedia, Airbnb, OpenAI, Google, Cohere, Snap, Getty Images, DeepSeek, Shopify. |
| Terms forbid score < 2.5 | Silent enforcers. Pinterest, Quora. | Aligned lax. Anthropic, Microsoft, Apple, xAI, Mistral, Stack Overflow, Salesforce/Slack, Nvidia, GitHub, Stability AI, Booking, Zoom, Atlassian, Common Crawl. |
37% of the sample writes anti-scraping into their Terms while leaving the machine-readable signal open. At these 13 companies, the declared position and the wire-level policy do not match.
Perplexity, LinkedIn, and Adobe are the strongest examples. Perplexity’s Terms score 6.5 on the forbid component. LinkedIn scores 6.5. Adobe scores 5.5. All three block zero AI training crawlers in their own robots.txt.
Pinterest and Quora reverse this. Both have permissive Terms (forbid 0) but their robots.txt blocks 10 of the 13 tracked AI training crawlers. The block is enforced at the CDN or WAF layer without any corresponding public statement in the Terms.
Companies that scrape after suing over scraping
LinkedIn spent years in court against hiQ Labs, a startup that scraped public LinkedIn profiles. The case reached the 9th Circuit in 2022. LinkedIn’s User Agreement forbids anyone else from using “crawlers, browser plugins and add-ons or any other technology” to access the platform. LinkedIn’s own Privacy Policy states that the company “may collect public information about you, such as professional-related news and accomplishments, and make it available as part of our Services.”
Getty Images filed suit against Stability AI in the UK High Court over training an image model on Getty’s library. The court ruled in 2025. Getty’s own Terms forbid “using any data mining, robots or similar data gathering or extraction methods.” The same behavior Getty accused Stability AI of is banned in Getty’s Terms for anyone else.
Reddit licensed its user content to Google for a reported $60 million a year, then sued Perplexity for accessing the same data without a contract. Reddit’s Terms grant Reddit “the right to use Your Content to train AI and machine learning models.”
X Corp’s liquidated damages clause
X Corp is the only company in the sample with a liquidated-damages clause. The clause specifies $15,000 in the United States and €15,000 in Europe per million posts scraped.

X’s Terms authorize the company to sue any party that “access or search or attempt to access or search the Services by any means (automated or otherwise) other than through our currently available, published interfaces.” At the same time, X’s own AI policy grants the company the right to use user content “for use with and training of our machine learning and artificial intelligence models.”
34 of 35 Tier-A companies leave enforcement to general remedies (injunction, actual damages), and only X Corp names a dollar amount.
Nine companies train on user content by default
Nine of the 35 Tier-A companies train AI models on their users’ content on an opt-out basis. Adobe, Anthropic, DeepSeek, GitHub, Microsoft, Mistral, OpenAI, Quora, and Stability AI all keep the behavior enabled by default. A user has to change the setting explicitly.
Anthropic’s Privacy Policy reads, “We may use your Inputs and Outputs to train and improve Anthropic AI models, unless you opt out through your account settings.” Microsoft’s Terms carry a parallel clause. “This data can help train our AI models in Microsoft Copilot unless you opt out.”
The scope varies by product. Adobe’s clause covers documents and images processed through its services. GitHub’s covers code and issues in public repositories (with different rules for private ones). Microsoft’s covers Copilot interactions and Microsoft 365 data. OpenAI’s covers ChatGPT conversations. Anthropic’s covers Claude conversations.
Clause language patterns across the industry
The 402 verified anti-scraping quotes cluster into a small number of legal templates. Companies are copying each other’s language.
The “robot, spider, scraper” template covers 17 sites. Netflix’s Terms include “use any robot, spider, scraper or other automated means to access the Netflix service.” Perplexity’s use “use any robot, spider, crawlers, scraper, or other automatic device, process, software or queries.” Expedia, Shopify, Snap, Airbnb, Discord, Unsplash, and Scribd all carry near-identical phrasings.
The “data mining, robots or similar data gathering” template covers 13 sites. Adobe’s clause reads “use any data mining or similar data gathering and extraction methods.” Cohere’s reads “use any data mining, robots, or similar data gathering or extraction methods.” Getty Images, IMDB, Forbes, Britannica, Duolingo, Goodreads, and Walmart share the same construction.
The AI-training-ban template covers 11 sites. Amazon’s Terms carry “You will not, and will not allow any third party to, use AI-generated content from the Amazon Services to, directly or indirectly, develop or improve large language or multimodal models.” Intuit’s Terms include “scraping, accessing, or downloading content that doesn’t belong to you, or training or developing artificial intelligence or machine learning models.” The Guardian, Time, Financial Times, and Unsplash use variations of the same clause.
The anti-circumvention template covers 12 sites. Netflix’s clause reads “circumvent, remove, alter, deactivate, degrade, block, obscure or thwart any of the content protections or other elements of the Netflix service.” Perplexity’s reads “circumvent, remove, alter, deactivate, degrade or thwart any technological measure or content protections of the Services.” Common Crawl, Getty, Time, Scribd, Slideshare, and Cisco carry near-identical language.
Terms of Service in this sample rely on shared legal templates. Individual companies vary the exact wording but the underlying construction is the same across each cluster. Anti-scraping language is adopted by default from industry norms, while extraction clauses are drafted separately for each company’s business model.
The rankings hold under alternative weights
The composite Index is the product of the forbid and take scores divided by ten. Because every scoring component is preserved as its own column, the same data can be re-weighted. Four alternative weightings produce the following top-5 lists.
| Scheme | Top 5 |
|---|---|
| Baseline (product) | Adobe, Perplexity, LinkedIn, Reddit, Yelp |
| Equal-weight sum | Adobe, Perplexity, Yelp, LinkedIn, Reddit |
| No AS5 (drop liquidated damages) | Adobe, Perplexity, LinkedIn, Reddit, Yelp |
| Forbid-only | Perplexity, LinkedIn, Adobe, X Corp, Midjourney |
| Take-only | GitHub, Adobe, Yelp, OpenAI, Google |
Heatmap:

Under the equal-weight sum, Yelp moves up two positions and GitHub jumps from 34th to 10th because its high take score (9) is no longer multiplied by a zero forbid. Dropping AS5 barely changes the ordering, since only X Corp has an AS5 clause.
Forbid-only isolates the sites with the strictest Terms. Perplexity leads at 6.5, LinkedIn at 6.5, Adobe at 5.5. Midjourney, Expedia, and Airbnb all move into the top-7 at forbid 4.5 each. Reddit, Amazon, and Yelp drop out because their high composite Index came from the take side.
Take-only produces the biggest ranking shifts. GitHub tops the list at take 9. Mistral moves from 28th to 6th. OpenAI and Google reach 4th and 5th.
Across all five weightings, Adobe is the only company in the top-5 under every scheme. Amazon, Perplexity, Reddit, and Yelp are in the top-10 under all five. The core of the leaderboard is robust. The rest of the ordering depends on whether the forbid or take side is weighted more heavily.
What the numbers add up to
Terms of Service and robots.txt are two independent signals about a site’s scraping stance. In this sample, they agree at the extremes and diverge in the middle.
At the strict end, 6 companies write both a strong Terms clause and a robots.txt block against AI training crawlers. X Corp, Reddit, Amazon, Yelp, Meta, and ByteDance.
At the permissive end, 14 companies write neither. Anthropic, Microsoft, Apple, and the rest score 0 on the forbid component and block no AI crawlers in robots.txt.
Between the two extremes, 15 companies split the signal. 13 declare a stance in Terms but do nothing in robots.txt (paper-only forbidders). 2 block in robots.txt without a matching Terms clause (silent enforcers).
The Terms themselves are mostly copies of shared legal templates. 17 sites use “robot, spider, scraper” phrasing. 13 use “data mining, robots or similar data gathering.” 12 share anti-circumvention wording. The double standard is largely the result of anti-scraping boilerplate being adopted by default from industry templates, while extraction clauses are drafted separately for each company’s business model.
The Index isolates the gap between those two sides. Under every alternative weighting, Adobe stays in the top-5. Amazon, Perplexity, Reddit, and Yelp stay in the top-10. The rest of the ranking depends on which side is weighted more heavily.
The data describes the gap between what a company’s own Terms forbid and what its Privacy Policy claims about data collected from elsewhere. It does not describe what a court would ultimately allow, only what the companies write in their own public documents.
Conclusion
The 402 verified clauses in this study describe what companies write, not what any court would ultimately enforce. Regulation in the US and EU is still working through the same questions. The July 2026 snapshot is a reference point for measuring how the gap changes as those decisions arrive.


