{"id":646991,"date":"2026-06-19T10:16:13","date_gmt":"2026-06-19T10:16:13","guid":{"rendered":"https:\/\/www.newsbeep.com\/uk\/646991\/"},"modified":"2026-06-19T10:16:13","modified_gmt":"2026-06-19T10:16:13","slug":"ais-biggest-casualty-could-be-history-itself","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/uk\/646991\/","title":{"rendered":"AI\u2019s biggest casualty could be history itself"},"content":{"rendered":"<p>Almost as soon as the web began, it started disappearing. The first ever <a href=\"https:\/\/www.independent.co.uk\/topic\/website\" rel=\"nofollow noopener\" target=\"_blank\">website<\/a>, which went live in August 1991, provided information on the <a href=\"https:\/\/www.independent.co.uk\/tech\/gen-z-nostalgia-retro-trends-dial-up-internet-b2836594.html\" title=\"Why Gen Z wants retro dial-up connections \u2013 and how the internet could be driven to extinction\" rel=\"nofollow noopener\" target=\"_blank\">world wide web project launched by Tim Berners-Lee<\/a>. But no one really knows what this page actually looked like. <\/p>\n<p>A screenshot of this first website was taken in 1992, but by then it had been updated countless times and was no longer an exact replica of the original. It wasn\u2019t until five years after the web was founded that efforts began to preserve its history, when computer engineer Brewster Kahle set up <a href=\"https:\/\/www.independent.co.uk\/arts-entertainment\/music\/news\/nirvana-1989-recording-aadam-jacobs-b2953719.html\" title=\"Volunteers turn a fan\u2019s recordings of 10,000 concerts into an online treasure trove\" rel=\"nofollow noopener\" target=\"_blank\">the Internet Archive<\/a> with the goal of providing \u201cuniversal access to all knowledge\u201d.<\/p>\n<p>It works by using software known as \u201ccrawlers\u201d to scour the <a href=\"https:\/\/www.independent.co.uk\/topic\/internet\" rel=\"nofollow noopener\" target=\"_blank\">internet<\/a> and take digital snapshots of public web pages. These are then stored and indexed in a massive <a href=\"https:\/\/www.independent.co.uk\/topic\/database\" rel=\"nofollow noopener\" target=\"_blank\">database<\/a> called the Wayback Machine that allows anyone to type in a URL address and see what it looked like at specific points in the past.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/06\/GettyImages-1426356126.jpeg\"  loading=\"lazy\" alt=\"Brewster Kahle set up the Internet Archive in 1992\" class=\"sc-1mc30lb-0 ggpMaE inline-gallery-btn\"\/>Brewster Kahle set up the Internet Archive in 1992 (Getty)<\/p>\n<p>In the three decades since it was founded, the <a href=\"https:\/\/www.independent.co.uk\/topic\/archive\" rel=\"nofollow noopener\" target=\"_blank\">archive<\/a> has made more than 1 trillion web captures. Some of these have been used to hold corporations and politicians accountable, provide evidence in court cases, and even prove that historical events happened. <\/p>\n<p>In July 2014, Russian-backed militants in Donetsk, Ukraine, posted on the social media site VKontakte that they had shot down a plane. When the plane turned out to be Malaysia Airlines Flight 17, and that all 298 people on board were killed, the post was swiftly deleted and the claim denied. But a record of the post remained on the Wayback Machine.<\/p>\n<p>In May 2020, Dominic Cummings, who was serving as the chief political adviser to prime minister Boris Johnson, claimed that he had warned of the threat of a coronavirus pandemic in a 2019 blog post. Within minutes of his speech, fact-checkers used the Wayback Machine to retrieve snapshots of his original blog, showing that he had updated it weeks after the lockdown had actually begun to make it look like he was prophetic. Once again, the tool was able to correct the historical record.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/06\/Internet-Archive-server-close-up-1536x991.jpg\"  loading=\"lazy\" alt=\"Some of the Internet Archive\u2019s servers \u2013 arguably the most prominent database of the online world\" class=\"sc-1mc30lb-0 ggpMaE inline-gallery-btn\"\/>Some of the Internet Archive\u2019s servers \u2013 arguably the most prominent database of the online world (Scott Beal\/Laughing Squid used under a Creative Commons license\/Nieman Lab)<\/p>\n<p>Yet despite its importance, the Internet Archive is now in peril. Online news outlets and sites like Reddit are increasingly blocking the crawlers used to gather snapshots \u2013 and AI is to blame.<\/p>\n<p>Organisations say the Internet Archive is just collateral damage in a battle between publishers and tech companies, which use similar crawling bots to scrape their copyrighted content to train their artificial intelligence models. Analysis by AI-detection startup Originality AI shows that 23 major news sites have already blocked crawlers used by the Wayback project, with the trend accelerating in recent months.<\/p>\n<p>Attempts are underway to save the Internet Archive, with non-profit Fight for the Future recently setting up a petition calling on media outlets to unblock the Wayback Machine\u2019s crawlers. <\/p>\n<p>Without [the Internet Archive], history becomes malleable and vulnerable to distortion \u2013 people and institutions can rewrite it with no accountability to facts<\/p>\n<p>\u201cIf anything, AI is the top reason why the Wayback Machine is more crucial than ever,\u201d the <a rel=\"nofollow noopener\" target=\"_blank\" href=\"https:\/\/clicks.independent.co.uk\/f\/a\/YBWMf9uSm0Xaj-sgltaIGw~~\/AAAHahA~\/8Xs0K4h5h6gp1en5GKjjRc9KBsl5lFPdO2r9UqFU0SOGK3l85YmlNWJTuXS9wU4goP9PhtXyiqBRU4SEArBHhTppO1dRTGnBH0Ez0WAO6LxBpQHUWKcK7F6sCDyZrmYsFkQhDpLxGc3-0oY4xDurd9TUqg9r-aKqZ1IxgWRthnR9VlxH9t63GIax8WObk7rZQ1fUblIFHWjzkys64SYXsl_efpUwIJ8jgKJnPX_GH9ROxHm1wfOTfRSaAwNmtt_8C4TBiZxi5vNKTyfQlwALuZBaxofsb0yTTq2-b3t2ceySVkIc5YONi1b7DYkeGzGqh_pSO1wTXMrpUR_-ytU6C846waUDocNu2Shtgs08_nYPoTCpN10m8Yng7JQt6IIgg0tUniboDTl_9XH13HpyoQ~~\">petition<\/a> states. \u201cCensorship and authoritarianism are growing, along with pressure to alter reporting and erase facts. The Wayback Machine makes every online news outlet it archives more resilient against pressure to remove stories that threaten the powerful.\u201d<\/p>\n<p>The web exists in a constant state of flux, where data is constantly overwritten, replaced, or lost to digital decay. Without a functional, universal archive, corporations and politicians can rewrite history without any accountability.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/06\/Internet_Archive_-_5079018246-copy.jpeg\"  loading=\"lazy\" alt=\"Online, IRL: the headquarters of the Internet Archive in San Francisco\" class=\"sc-1mc30lb-0 ggpMaE inline-gallery-btn\"\/>Online, IRL: the headquarters of the Internet Archive in San Francisco (Beatrice Murch\/CC BY 2.0)<\/p>\n<p>The term \u201cOrwellian\u201d is often overused, but the Internet Archive provides an actual antidote to Winston Smith\u2019s position in the Ministry of Truth, where the protagonist of 1984 was tasked with rewriting and erasing historical documents using a \u201cmemory hole\u201d device. The Archive is integral to avoiding this kind of dystopia.<\/p>\n<p>\u201cThe average life of a web page is 100 days before it\u2019s changed or deleted,\u201d Kahle, the founder of the Internet Archive, said on a recent episode of the podcast Close All Tabs. \u201cIf we do not actively collect them and preserve them and keep them accessible, we are living in the memory hole universe of George Orwell.\u201d<\/p>\n<p>The work of Kahle and those at the Internet Archive is preserving this digital history, which is increasingly part of global history. Without it, history becomes malleable and vulnerable to distortion \u2013 people and institutions can rewrite it with no accountability to facts.<\/p>\n<p>AI is doing many harms to the internet: Supercharging the spread of misinformation, accelerating cyber attacks, and flooding the web with slop. But this might be the most pernicious. If we lose the Internet Archive, it won\u2019t just be the first-ever web page lost to history.<\/p>\n","protected":false},"excerpt":{"rendered":"Almost as soon as the web began, it started disappearing. The first ever website, which went live in&hellip;\n","protected":false},"author":2,"featured_media":646992,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[554,733,4308,86,56,54,55],"class_list":["post-646991","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-technology","tag-uk","tag-united-kingdom","tag-unitedkingdom"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/646991","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/comments?post=646991"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/646991\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media\/646992"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media?parent=646991"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/categories?post=646991"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/tags?post=646991"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}