{"id":44074,"date":"2025-08-04T21:09:11","date_gmt":"2025-08-04T21:09:11","guid":{"rendered":"https:\/\/www.newsbeep.com\/uk\/44074\/"},"modified":"2025-08-04T21:09:11","modified_gmt":"2025-08-04T21:09:11","slug":"perplexity-accused-of-scraping-websites-that-explicitly-blocked-ai-scraping","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/uk\/44074\/","title":{"rendered":"Perplexity accused of scraping websites that explicitly blocked AI scraping"},"content":{"rendered":"<p id=\"speakable-summary\" class=\"wp-block-paragraph\">AI startup Perplexity is crawling and scraping content from websites that have explicitly indicated they don\u2019t want to be scraped, according to internet infrastructure provider Cloudflare.<\/p>\n<p class=\"wp-block-paragraph\">On Monday, Cloudflare <a href=\"https:\/\/blog.cloudflare.com\/perplexity-is-using-stealth-undeclared-crawlers-to-evade-website-no-crawl-directives\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">published research<\/a> saying it observed the AI startup ignore blocks and hide its crawling and scraping activities. The network infrastructure giant accused Perplexity of obscuring its identity when trying to scrape web pages \u201cin an attempt to circumvent the website\u2019s preferences,\u201d Cloudflare\u2019s researchers wrote.<\/p>\n<p class=\"wp-block-paragraph\">AI products like those offered by Perplexity rely on gobbling up large amounts of data from the internet, and AI startups have long scraped text, images, and videos from the internet many times without permission to make their products work. In recent times, websites have tried to fight back by using the web standard Robots.txt file, which tells search engines and AI companies which pages can be indexed and which shouldn\u2019t, efforts <a href=\"https:\/\/www.reuters.com\/technology\/artificial-intelligence\/multiple-ai-companies-bypassing-web-standard-scrape-publisher-sites-licensing-2024-06-21\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">that have seen mixed results so far<\/a>.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Perplexity appears to be willingly circumventing these blocks by changing its bots\u2019 \u201cuser agent,\u201d meaning a signal that identifies a website visitor by their device and version type, as well as changing their autonomous system networks, or ASN, essentially a number that identifies large networks on the internet, according to Cloudflare.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cThis activity was observed across tens of thousands of domains and millions of requests per day. We were able to fingerprint this crawler using a combination of machine learning and network signals,\u201d read Cloudflare\u2019s post.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Perplexity spokesperson Jesse Dwyer dismissed Cloudflare\u2019s blog post as a \u201csales pitch,\u201d adding in an email to TechCrunch that the screenshots in the post \u201cshow that no content was accessed.\u201d In a follow-up email, Dwyer claimed the bot named in the Cloudflare blog \u201cisn\u2019t even ours.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Cloudflare said it first noticed the behavior after its customers complained that Perplexity was crawling and scraping their sites, even after they added rules on their Robots file and for specifically blocking Perplexity\u2019s known bots. Cloudflare said it then performed tests to check and confirmed that Perplexity was circumventing these blocks.\u00a0<\/p>\n<p>Techcrunch event<\/p>\n<p>\n\t\t\t\t\t\t\t\t\tSan Francisco<br \/>\n\t\t\t\t\t\t\t\t\t\t\t\t\t|<br \/>\n\t\t\t\t\t\t\t\t\t\t\t\t\tOctober 27-29, 2025\n\t\t\t\t\t\t\t<\/p>\n<p class=\"wp-block-paragraph\">\u201cWe observed that Perplexity uses not only their declared user-agent, but also a generic browser intended to impersonate Google Chrome on macOS when their declared crawler was blocked,\u201d according to Cloudflare.\u00a0\u00a0<\/p>\n<p class=\"wp-block-paragraph\">The company also said that it has de-listed Perplexity\u2019s bots from its verified list and added new techniques to block them.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Cloudflare has recently taken a public stance against AI crawlers. Last month, Cloudflare <a href=\"https:\/\/techcrunch.com\/2025\/07\/01\/cloudflare-launches-a-marketplace-that-lets-websites-charge-ai-bots-for-scraping\/\" rel=\"nofollow noopener\" target=\"_blank\">announced the launch of a marketplace<\/a> allowing website owners and publishers to charge AI scrapers who visit their sites. Cloudflare\u2019s chief executive Matthew Prince <a href=\"https:\/\/www.cfr.org\/event\/bernard-l-schwartz-annual-lecture-matthew-prince-cloudflare\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">sounded the alarm<\/a> at the time, saying AI is breaking the business model of the internet, particularly publishers. Last year, Cloudflare also <a href=\"https:\/\/techcrunch.com\/2024\/07\/03\/cloudflare-launches-a-tool-to-combat-ai-bots\/#:~:text=Cloudflare%20has%20set%20up%20a,demand%20for%20model%20training%20data.\" rel=\"nofollow noopener\" target=\"_blank\">launched a free tool<\/a> to prevent bots from scraping websites to train AI.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">This is not the first time Perplexity is accused of scraping without authorization.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Last year, news outlets, <a href=\"https:\/\/www.wired.com\/story\/perplexity-is-a-bullshit-machine\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">such as Wired<\/a>, alleged <a href=\"http:\/\/techcrunch.com\/2024\/07\/02\/news-outlets-are-accusing-perplexity-of-plagiarism-and-unethical-web-scraping\/\" rel=\"nofollow noopener\" target=\"_blank\">Perplexity was plagiarizing their content<\/a>. Weeks later, Perplexity\u2019s CEO Aravind Srinivas <a href=\"http:\/\/techcrunch.com\/2024\/10\/30\/perplexitys-ceo-punts-on-defining-plagiarism\/\" rel=\"nofollow noopener\" target=\"_blank\">was unable to immediately answer<\/a> when asked to provide the company\u2019s definition of plagiarism during an interview with TechCrunch\u2019s Devin Coldewey at the Disrupt 2024 conference.<\/p>\n","protected":false},"excerpt":{"rendered":"AI startup Perplexity is crawling and scraping content from websites that have explicitly indicated they don\u2019t want to&hellip;\n","protected":false},"author":2,"featured_media":44075,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[554,733,11624,4308,25178,11173,1923,2991,25179,86,56,54,55],"class_list":["post-44074","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificial-intelligence-ai","tag-artificialintelligence","tag-bots","tag-cloudflare","tag-llms","tag-perplexity","tag-scraping","tag-technology","tag-uk","tag-united-kingdom","tag-unitedkingdom"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/44074","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/comments?post=44074"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/44074\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media\/44075"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media?parent=44074"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/categories?post=44074"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/tags?post=44074"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}