{"id":635284,"date":"2026-04-29T00:53:30","date_gmt":"2026-04-29T00:53:30","guid":{"rendered":"https:\/\/www.newsbeep.com\/ca\/635284\/"},"modified":"2026-04-29T00:53:30","modified_gmt":"2026-04-29T00:53:30","slug":"talkie-is-a-vintage-llm-trained-on-pre-1930-data-to-help-facilitate-time-travel","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/ca\/635284\/","title":{"rendered":"Talkie Is a &#8216;Vintage LLM&#8217; Trained on Pre-1930 Data to Help Facilitate &#8216;Time Travel&#8217;"},"content":{"rendered":"<p>If you\u2019ve ever heard the term \u201cvintage LLM\u201d, you might have found yourself wondering if the AI-pocalypse has really been going on for long enough that early chatbots are worthy of nostalgia. Happily, though, that\u2019s not what the term means; instead, it applies to an LLM that seeks to emulate the perspective of a certain point in the past.<\/p>\n<p>The idea is that you restrict the training data provided to the model to material published before a given date. In the case of <a href=\"https:\/\/talkie-lm.com\/introducing-talkie\" rel=\"nofollow noopener\" target=\"_blank\">Talkie<\/a>, aka 13B 1930 LM, the cutoff is, as the name suggests, the year 1930. This choice of year might seem arbitrary, but it\u2019s not: as we <a href=\"https:\/\/gizmodo.com\/wikiflix-helps-you-catch-up-on-films-that-just-entered-the-public-domain-2000708070\" rel=\"nofollow noopener\" target=\"_blank\">discussed here back in January<\/a>, many forms of copyright expire on January 1 of the year that comes 95 years after the copyrighted material was released. This means that a whole lot of material released in 1930 went into the public domain at the start of this year.<\/p>\n<p>Choosing 1930 as a cut-off date thus allows Talkie to sidestep the question of how to navigate copyright, an issue that has been blithely ignored by a problem for other LLMs. Which, OK, this is all interesting, but what is it for?<\/p>\n<p>In answering that question, Talkie\u2019s creators lean heavily on two sources: a <a href=\"https:\/\/owainevans.github.io\/talk-transcript.html\" rel=\"nofollow noopener\" target=\"_blank\">talk<\/a> given by AI researcher <a href=\"https:\/\/owainevans.github.io\/\" rel=\"nofollow noopener\" target=\"_blank\">Owain Evans<\/a>, from which the term \u201cvintage LLM\u201d originated, and a <a href=\"https:\/\/www.calcifercomputing.com\/reports\/tlm\" rel=\"nofollow noopener\" target=\"_blank\">paper<\/a> on \u201ctemporal language models\u201d originating from a company called Calcifer Computing, whose <a href=\"https:\/\/www.calcifercomputing.com\/products\" rel=\"nofollow noopener\" target=\"_blank\">business<\/a> apparently involves providing \u201cnon-recurring engineering services for clients with problems interesting enough to warrant our attention.\u201d<\/p>\n<p>Evans\u2019 talk couches its ideas in the sort of hyperbole that seems compulsory for AI proponents: \u201cThe first humanistic motivation [for vintage LLMs] is time travel,\u201d he explains modestly. \u201cWhat would it be like to communicate with someone from 1700?\u201d To answer this question, he proposes the idea of models trained on data that cuts off at a certain time\u2014exactly the sort of thing Talkie is doing, in other words.<\/p>\n<p>The Calcifer Computing paper is less violet-hued, discussing the challenge of how LLMs can account for the way that aspects of language\u2014words\u2019 meanings, speech patterns, vocabulary\u2014change over time. This is genuinely interesting, and apparently inspired one of the first uses for Talkie, which was to provide a subjective rating of the \u201csurprisingness\u201d of various post-1930 events.<\/p>\n<p>The really interesting question, though, is how reliably an LLM trained on data that cuts off at a certain date can predict what will happen after that date. This feels like a scaled-down sociological analogue of the more fundamental question of determinism, which asks whether knowing everything about a system\u2019s initial state allows you to predict that system\u2019s future states.<\/p>\n<p>Of course, you can feed an LLM information until the cows come home; it will never and could never know everything possible about the state of the world in 1930, or 69 BC, or 5:30pm yesterday. But still, the idea of giving an LLM a solid grounding in history along with a large amount of information about the state of the world in 1930, and then asking, \u201cWhat happens next?\u201d\u2026 even for someone generally skeptical of LLMs (like, y\u2019know, me), that\u2019s an interesting question. Being AI people, Talkie\u2019s creators aren\u2019t satisfied with predicting the future; they also allude to a question posed by Google DeepMind CEO Demis Hassabis, who once asked whether an LLM trained with data that cuts off at 1911 could discover general relativity.<\/p>\n<p>Sadly, it doesn\u2019t appear that there\u2019s an answer to either question yet. The rest of the paper is devoted mostly to explaining the various challenges of getting Talkie to work reliably, foremost among them the lack of reliable training data. Talkie is trained on data scanned from physical sources, making reliable character recognition facilities extremely important. There\u2019s also the challenge of what the authors call \u201ccontamination,\u201d i.e., the leakage of post-1930s material into the training data.<\/p>\n<p>At this point, Talkie falls into the \u201cpotentially interesting and apparently harmless\u201d category of LLM\/AI agent projects, which, honestly, these days feels like about the best we can hope for. The project\u2019s site contains a live feed of the LLM answering questions posed by \u2026 another LLM. As this post was being written, Talkie was describing an 1882 cricket match, and if you\u2019ll bear with me here:<\/p>\n<p> <img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-2000751759\" src=\"https:\/\/www.newsbeep.com\/ca\/wp-content\/uploads\/2026\/04\/Screenshot-2026-04-28-at-12.20.18.jpg\" alt=\"Talkie talks about cricket\" width=\"1330\" height=\"1098\"  \/>\u00a9 Talkie <\/p>\n<p>This all seems very evocative, but unfortunately, I am a cricket nerd, and I can assure you that the match Talkie is describing never happened. The only Test played between Australia and England in 1882 was <a href=\"https:\/\/www.espncricinfo.com\/series\/australia-tour-of-england-1882-61352\/england-vs-australia-only-test-62404\/full-scorecard\" rel=\"nofollow noopener\" target=\"_blank\">at the Oval<\/a> in August that year, and it was perhaps the most famous ever played\u2014the match was such a disastrous defeat for England that one London paper penned an obituary for English cricket, giving birth to The Ashes, a biennial series between the two countries that continues to this day.<\/p>\n<p>It seems a curious choice for Talkie to describe a fictional match, and to place it specifically in the year of such a famous real match. It seems like a good guess that the training data is particularly well-stocked for descriptions of Test matches in that year, so out of interest, I asked Talkie to describe the actual Ashes test of 1882. Sadly, the result was even less accurate, featuring incorrect scores and at least one completely made-up player. Still, the descriptions it produces are certainly colorful and believable, so that\u2019s \u2026 something?<\/p>\n<p>Anyway, if Talkie does manage to keep its feet planted in reality and predict World War II, we look forward to letting you know about its verdicts on famous events like the siege of Leningrad, the D-Day landings at Brittany, and, of course, the surprise Japanese attack on Diego Garcia. Can it capture the unique timbre of the speeches of UK Prime Minister Winfield Cromwell or the terrifying machine-gun delivery of German dictator Rudolf Schei\u00dfe? That remains to be seen.<\/p>\n","protected":false},"excerpt":{"rendered":"If you\u2019ve ever heard the term \u201cvintage LLM\u201d, you might have found yourself wondering if the AI-pocalypse has&hellip;\n","protected":false},"author":2,"featured_media":635285,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[62,276,277,49,48,237467,61,237468],"class_list":["post-635284","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-ca","tag-canada","tag-talkie","tag-technology","tag-vintage-llms"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/posts\/635284","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/comments?post=635284"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/posts\/635284\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/media\/635285"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/media?parent=635284"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/categories?post=635284"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/tags?post=635284"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}