Of the ten highest-scoring companies on the AI Hypocrisy Index, five block zero AI training crawlers in their robots.txt. The top three, Adobe, Perplexity, and LinkedIn, are all in that group. Their Terms score high on the anti-scraping component. Their robots.txt is open to every AI crawler in the sample. The declared rule and the enforced rule disagree at the top of the list.
That’s one of ten findings from the July 2026 baseline. We read the Terms of Service and robots.txt of 35 leading AI companies and platforms, plus the 150 most-visited consumer websites, and classified 402 clauses with a verbatim quote for every one. 19 of 35 AI companies explicitly ban scraping in their Terms while claiming rights to user or public data. 85 of the 150 top consumer sites do the same. 47 instances of one AI platform’s robots.txt disallowing a rival AI company’s crawler. X Corp charges $15,000 (and €15,000) per million posts scraped, the only company in the sample with a liquidated-damages clause. Nine companies train on user content by default, opt-out. Adobe tops the Index at 4.4 out of 10 and stays there under every alternative weighting scheme we tried. Companies copy each other’s legal language. 17 firms share nearly identical “robot, spider, scraper” phrasing across their Terms.
This is the July 2026 snapshot. Every claim is backed by an exact-substring-verified quote from the source document.
What we measured
Two tiers. Tier A is 35 companies that build AI models or operate crawlers. For each Tier A company we read the Terms of Service, Privacy Policy, and robots.txt. Tier B is the 150 most-visited consumer websites drawn from the Tranco top ranking (list 94Y82). For Tier B we read Terms of Service and robots.txt only. HasData and its own domain were excluded from the sample.
Documents were fetched on 23 July 2026 with a custom User-Agent (DoubleStandardIndexResearch/1.0), at one request per second per domain. No paywalls, logins, or bot protections were circumvented.
Each document was classified against a 12-clause taxonomy. The taxonomy has 7 anti-scraping clauses and 5 extraction clauses. An LLM classified each document at temperature zero, returning strict JSON. Every reported clause carries a verbatim quotation that is programmatically verified as an exact substring of the source document. Classification was run twice. 402 findings passed verification at a 100% exact-substring rate. 12 findings failed verification and were dropped.
Each Tier-A company gets two scores on a 0 to 10 scale. The forbid score counts anti-scraping clauses. The take score counts data extraction. The AI Hypocrisy Index is their product divided by ten. A pure blocker or a pure scraper stays low. The Index is high only when a company scores high on both sides.
The numbers at a glance:
| Metric | Tier A (35 companies) | Tier B (150 top sites) |
|---|---|---|
| Bans scraping in Terms | 19 | 85 |
| Requires prior written consent | 6 | 11 |
| Explicit AI-training ban on their own content | 5 | 22 |
| Blocks 1+ AI training crawlers in robots.txt | 8 | 44 |
Tier-A-only findings, since Privacy Policies were not read for Tier B. 9 companies train on user content by default (opt-out). 7 companies claim rights to “publicly available” web data. 1 company has a liquidated-damages clause (X Corp).
The measurements use these definitions.
- Anti-scraping ban. Terms of Service contain at least one of the 7 AS clauses (general scraping ban, absolute ban, prior written consent, AI-training ban on their content, liquidated damages, anti-circumvention, API-only access).
- AI training crawler. A user-agent operated by an AI vendor for training-data collection. 13 were tracked, including GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, PerplexityBot, CCBot, and Meta-ExternalAgent.
- Paper-only forbidder. Terms declare an anti-scraping stance (forbid score ≥ 2.5) and robots.txt blocks zero of the 13 tracked AI training crawlers.
- Silent enforcer. Terms are permissive (forbid score < 2.5) and robots.txt blocks at least one AI training crawler.
- Verbatim verification. A reported clause is included only if its exact wording is found as an exact substring of the fetched source document.
One of 35 Tier-A Terms pages was a JavaScript-only site and could not be read in text form. The historical (Wayback) dimension is partial and is used for context only.
The AI Hypocrisy Index leaderboard
Adobe tops the Index at 4.4 out of 10. Perplexity is second at 3.9. LinkedIn, Reddit, and Yelp round out the top five between 2.4 and 2.6.

Table 1 lists the top 15 Tier-A companies with their forbid score, take score, composite Index, and the number of AI training crawlers each blocks in its own robots.txt.
| # | Company | Domain | Forbid | Take | Index | AI bots blocked |
|---|---|---|---|---|---|---|
| 1 | Adobe | adobe.com | 5.5 | 8 | 4.4 | 0 |
| 2 | Perplexity | perplexity.ai | 6.5 | 6 | 3.9 | 0 |
| 3 | linkedin.com | 6.5 | 4 | 2.6 | 0 | |
| 4 | reddit.com | 4 | 6 | 2.4 | 11 | |
| 5 | Yelp | yelp.com | 3 | 8 | 2.4 | 11 |
| 6 | Amazon | amazon.com | 3.5 | 6 | 2.1 | 9 |
| 7 | X Corp | x.com | 5.0 | 4 | 2.0 | 10 |
| 8 | OpenAI | openai.com | 2.5 | 7 | 1.75 | 0 |
| 9 | google.com | 2.5 | 7 | 1.75 | 0 | |
| 10 | Meta | facebook.com | 2.5 | 6 | 1.5 | 11 |
| 11 | Cohere | cohere.com | 2.5 | 6 | 1.5 | 0 |
| 12 | Snap | snap.com | 2.5 | 6 | 1.5 | 0 |
| 13 | Midjourney | midjourney.com | 4.5 | 3 | 1.35 | 0 |
| 14 | Expedia | expedia.com | 4.5 | 3 | 1.35 | 0 |
| 15 | Airbnb | airbnb.com | 4.5 | 3 | 1.35 | 0 |
The Index is high only when both scores are high. Anthropic takes (score 5) but has permissive Terms (forbid 0), so its composite Index is 0.0. GitHub takes even more broadly (score 9) but writes nothing anti-scraping into its Terms, so its Index is also 0.0. The Index measures the gap between what a company writes in its Terms and what it does with data. A high score on one side alone does not put a company on the leaderboard.
The forbid-vs-take split
The Index has two axes. Plotted against each other, they separate two distinct categories of double standard.

The upper-left quadrant contains the pure AI builders. OpenAI and Google both score 2.5 on forbid and 7 on take. Their Terms are moderate on anti-scraping. Both claim rights to “publicly available” data and both operate a training crawler.
The upper-right quadrant contains the content platforms and creative-software firms. Adobe (5.5 forbid, 8 take), Perplexity (6.5, 6), Reddit (4, 6), and Yelp (3, 8) all restrict access to their own content while claiming broad rights over user content. This is where the Index scores highest.
X Corp and LinkedIn are in the mid-right. Both score high on forbid (5.0 and 6.5) but only 4 on take. The take scores reflect narrower user-content licensing than the platforms in the top-right.
The bottom-left is where forbid is 0 and take ranges from 2 to 5. Anthropic, Microsoft, Apple, Nvidia, Salesforce/Slack, xAI, Mistral, Stability AI, and Stack Overflow are all here. These companies score positive on take but not on forbid, which puts their composite Index at 0.0.
Banning scraping is the norm
19 of the 35 AI companies write an explicit anti-scraping clause into their Terms of Service. Across the 150 most-visited consumer websites, 85 do the same. It’s the default across the top sites in the sample.

The AI builders themselves are in this group. Google’s Terms prohibit “scraping content that doesn’t belong to you.” OpenAI bans “automatically or programmatically extracting data or Output.” Both companies also describe collecting “information that is publicly available on the internet” to develop their models, in policy documents on the same domain.
Seven of the 35 Tier-A companies write both sides of that contradiction into their published documents. Adobe, Booking, Getty Images, Google, LinkedIn, OpenAI, and Yelp each carry an anti-scraping clause in one document and a claim of rights over “publicly available” web data in another.
AI labs block each other in robots.txt
Reading each AI platform’s robots.txt against every other AI company’s crawler user-agent, we found 47 instances of a platform disallowing a rival AI training crawler. Reddit and Yelp each block 11 of the 13 tracked AI training bots. Meta blocks 11. X Corp, Pinterest, and Quora each block 10. Amazon blocks 9. ByteDance/TikTok blocks 8.

Which platform blocks which rival is visible in the matrix. X Corp’s robots.txt disallows OpenAI, Anthropic, Google, Perplexity, and six others. Meta disallows ten rivals. The refusal is mutual across almost every pair of platforms in the sample.

For an AI company crawling the web, the pool of allowed sources is narrower than the phrase “publicly available” suggests. The columns of the matrix are the AI vendors themselves. OpenAI, Anthropic, Google, Perplexity, Apple, Amazon, Meta, ByteDance, Mistral, Cohere, Common Crawl, and Microsoft all appear on the “blocked” side of at least one platform in this sample.
Where robots.txt disagrees with Terms
The Terms of Service and the robots.txt are separate signals from the same company. In this sample, they disagree in both directions.
| robots.txt blocks 1+ AI training crawler | robots.txt blocks 0 AI training crawlers | |
|---|---|---|
| Terms forbid score ≥ 2.5 | Aligned strict. X Corp, Reddit, Amazon, Yelp, Meta, TikTok. | Paper-only forbidders. Adobe, Perplexity, LinkedIn, Midjourney, Expedia, Airbnb, OpenAI, Google, Cohere, Snap, Getty Images, DeepSeek, Shopify. |
| Terms forbid score < 2.5 | Silent enforcers. Pinterest, Quora. | Aligned lax. Anthropic, Microsoft, Apple, xAI, Mistral, Stack Overflow, Salesforce/Slack, Nvidia, GitHub, Stability AI, Booking, Zoom, Atlassian, Common Crawl. |
37% of the sample writes anti-scraping into their Terms while leaving the machine-readable signal open. At these 13 companies, the declared position and the wire-level policy do not match.
Perplexity, LinkedIn, and Adobe are the strongest examples. Perplexity’s Terms score 6.5 on the forbid component. LinkedIn scores 6.5. Adobe scores 5.5. All three block zero AI training crawlers in their own robots.txt.
Pinterest and Quora reverse this. Both have permissive Terms (forbid 0) but their robots.txt blocks 10 of the 13 tracked AI training crawlers. The block is enforced at the CDN or WAF layer without any corresponding public statement in the Terms.
Companies that scrape after suing over scraping
LinkedIn spent years in court against hiQ Labs, a startup that scraped public LinkedIn profiles. The case reached the 9th Circuit in 2022. LinkedIn’s User Agreement forbids anyone else from using “crawlers, browser plugins and add-ons or any other technology” to access the platform. LinkedIn’s own Privacy Policy states that the company “may collect public information about you, such as professional-related news and accomplishments, and make it available as part of our Services.”
Getty Images filed suit against Stability AI in the UK High Court over training an image model on Getty’s library. The court ruled in 2025. Getty’s own Terms forbid “using any data mining, robots or similar data gathering or extraction methods.” The same behavior Getty accused Stability AI of is banned in Getty’s Terms for anyone else.
Reddit did not go to court, but its handling of user content points the same way. Reddit’s Terms grant Reddit “the right to use Your Content to train AI and machine learning models.” Reddit licensed that user content to Google for a reported $60 million a year.
X Corp’s liquidated damages clause
X Corp is the only company in the sample with a liquidated-damages clause. The clause specifies $15,000 in the United States and €15,000 in Europe per million posts scraped.

X’s Terms authorize the company to sue any party that “access or search or attempt to access or search the Services by any means (automated or otherwise) other than through our currently available, published interfaces.” At the same time, X’s own AI policy grants the company the right to use user content “for use with and training of our machine learning and artificial intelligence models.”
34 of 35 Tier-A companies leave enforcement to general remedies (injunction, actual damages), and only X Corp names a dollar amount.
Nine companies train on user content by default
Nine of the 35 Tier-A companies train AI models on their users’ content on an opt-out basis. Adobe, Anthropic, DeepSeek, GitHub, Microsoft, Mistral, OpenAI, Quora, and Stability AI all keep the behavior enabled by default. A user has to change the setting explicitly.
Anthropic’s Privacy Policy reads, “We may use your Inputs and Outputs to train and improve Anthropic AI models, unless you opt out through your account settings.” Microsoft’s Terms carry a parallel clause. “This data can help train our AI models in Microsoft Copilot unless you opt out.”
The scope varies by product. Adobe’s clause covers documents and images processed through its services. GitHub’s covers code and issues in public repositories (with different rules for private ones). Microsoft’s covers Copilot interactions and Microsoft 365 data. OpenAI’s covers ChatGPT conversations. Anthropic’s covers Claude conversations.
Clause language patterns across the industry
The 402 verified anti-scraping quotes cluster into a small number of legal templates. Companies are copying each other’s language.
The “robot, spider, scraper” template covers 17 sites. Netflix’s Terms include “use any robot, spider, scraper or other automated means to access the Netflix service.” Perplexity’s use “use any robot, spider, crawlers, scraper, or other automatic device, process, software or queries.” Expedia, Shopify, Snap, Airbnb, Discord, Unsplash, and Scribd all carry near-identical phrasings.
The “data mining, robots or similar data gathering” template covers 13 sites. Adobe’s clause reads “use any data mining or similar data gathering and extraction methods.” Cohere’s reads “use any data mining, robots, or similar data gathering or extraction methods.” Getty Images, IMDB, Forbes, Britannica, Duolingo, Goodreads, and Walmart share the same construction.
The AI-training-ban template covers 11 sites. Amazon’s Terms carry “You will not, and will not allow any third party to, use AI-generated content from the Amazon Services to, directly or indirectly, develop or improve large language or multimodal models.” Intuit’s Terms include “scraping, accessing, or downloading content that doesn’t belong to you, or training or developing artificial intelligence or machine learning models.” The Guardian, Time, Financial Times, and Unsplash use variations of the same clause.
The anti-circumvention template covers 12 sites. Netflix’s clause reads “circumvent, remove, alter, deactivate, degrade, block, obscure or thwart any of the content protections or other elements of the Netflix service.” Perplexity’s reads “circumvent, remove, alter, deactivate, degrade or thwart any technological measure or content protections of the Services.” Common Crawl, Getty, Time, Scribd, Slideshare, and Cisco carry near-identical language.
Terms of Service in this sample rely on shared legal templates. Individual companies vary the exact wording but the underlying construction is the same across each cluster. Anti-scraping language is adopted by default from industry norms, while extraction clauses are drafted separately for each company’s business model.
The rankings hold under alternative weights
The composite Index is the product of the forbid and take scores divided by ten. Because every scoring component is preserved as its own column, the same data can be re-weighted. Four alternative weightings produce the following top-5 lists.
| Scheme | Top 5 |
|---|---|
| Baseline (product) | Adobe, Perplexity, LinkedIn, Reddit, Yelp |
| Equal-weight sum | Adobe, Perplexity, Yelp, LinkedIn, Reddit |
| No AS5 (drop liquidated damages) | Adobe, Perplexity, LinkedIn, Reddit, Yelp |
| Forbid-only | Perplexity, LinkedIn, Adobe, X Corp, Midjourney |
| Take-only | GitHub, Adobe, Yelp, OpenAI, Google |
Heatmap:

Under the equal-weight sum, Yelp moves up two positions and GitHub jumps from 34th to 10th because its high take score (9) is no longer multiplied by a zero forbid. Dropping AS5 barely changes the ordering, since only X Corp has an AS5 clause.
Forbid-only isolates the sites with the strictest Terms. Perplexity leads at 6.5, LinkedIn at 6.5, Adobe at 5.5. Midjourney, Expedia, and Airbnb all move into the top-7 at forbid 4.5 each. Reddit, Amazon, and Yelp drop out because their high composite Index came from the take side.
Take-only produces the biggest ranking shifts. GitHub tops the list at take 9. Mistral moves from 28th to 6th. OpenAI and Google reach 4th and 5th.
Across all five weightings, Adobe is the only company in the top-5 under every scheme. Amazon, Perplexity, Reddit, and Yelp are in the top-10 under all five. The core of the leaderboard is robust. The rest of the ordering depends on whether the forbid or take side is weighted more heavily.
What the numbers add up to
Terms of Service and robots.txt are two independent signals about a site’s scraping stance. In this sample, they agree at the extremes and diverge in the middle.
At the strict end, 6 companies write both a strong Terms clause and a robots.txt block against AI training crawlers. X Corp, Reddit, Amazon, Yelp, Meta, and ByteDance.
At the permissive end, 14 companies write neither. Anthropic, Microsoft, Apple, and the rest score 0 on the forbid component and block no AI crawlers in robots.txt.
Between the two extremes, 15 companies split the signal. 13 declare a stance in Terms but do nothing in robots.txt (paper-only forbidders). 2 block in robots.txt without a matching Terms clause (silent enforcers).
The Terms themselves are mostly copies of shared legal templates. 17 sites use “robot, spider, scraper” phrasing. 13 use “data mining, robots or similar data gathering.” 12 share anti-circumvention wording. The double standard is largely the result of anti-scraping boilerplate being adopted by default from industry templates, while extraction clauses are drafted separately for each company’s business model.
The Index isolates the gap between those two sides. Under every alternative weighting, Adobe stays in the top-5. Amazon, Perplexity, Reddit, and Yelp stay in the top-10. The rest of the ranking depends on which side is weighted more heavily.
The data describes the gap between what a company’s own Terms forbid and what its Privacy Policy claims about data collected from elsewhere. It does not describe what a court would ultimately allow, only what the companies write in their own public documents.
Conclusion
The 402 verified clauses in this study describe what companies write, not what any court would ultimately enforce. Regulation in the US and EU is still working through the same questions. The July 2026 snapshot is a reference point for measuring how the gap changes as those decisions arrive.


