Help support TMP


"Vintage language models" Topic


4 Posts

All members in good standing are free to post here. Opinions expressed here are solely those of the posters, and have not been cleared with nor are they endorsed by The Miniatures Page.

In order to respect possible copyright issues, when quoting from a book or article, please quote no more than three paragraphs.

For more information, see the TMP FAQ.


Back to the Artificial Intelligence (AI) Message Board


Areas of Interest

General

Featured Hobby News Article


Featured Link


Featured Ruleset


Featured Showcase Article

Little Yellow Clamps

Need some low-pressure clamps?


Featured Profile Article

Those Blasted Trees

How do you depict "shattered forest" on the tabletop?


Featured Book Review


306 hits since 26 Jul 2026
©1994-2026 Bill Armintrout
Comments or corrections?

pellen26 Jul 2026 3:03 p.m. PST

Thought some here may enjoy this. I posted about it somewhere on BGG a few months ago.

Talkie 1930 is a "vintage language model". Only texts from before 1931 were used to train it (give or take some small amount of newer texts that were accidentally included… they discuss that a bit in the linked page). Seems like it did not get much mainstream attention when it was released ~3 months ago:

link

It's a pretty small for a "large" language model, not even on the level of the first public versions of ChatGPT (if any of you remember that; very primitive compared to what we have now).

Of course you can instruct any LLM to role-play and pretend that it's 1930 and it doesn't know the future, but the point of a vintage language model is that it really can't leak any modern information or bias into its answers.

There have been a few other vintage language models, and no doubt more will be released. Talkie 1930 seems to be the most useful one available for now. Other models have different cut-off dates. Only using texts from before 1900 seems like a popular one. But from what I understand there is a hard limit somewhere in the late 19th century. Before that date there just isn't enough text available to be able to train a large language model from scratch.

TheBeast Supporting Member of TMP27 Jul 2026 8:59 a.m. PST

Given the number of things that 'haven't aged well', the further back you go, the scarier it sounds.

I was born at the beginning of the '50s, and the baggage I carry is impressive to me.

Doug

Personal logo Parzival Supporting Member of TMP27 Jul 2026 3:41 p.m. PST

Yeah, I'm kind of wondering what "texts" are being used as references; there's a lot of garbage ideas floating around in human history. But then, there's plenty of garbage ideas floating around the Internet right now. Either way, GI,GO.

pellen28 Jul 2026 4:00 a.m. PST

Mostly books and newspapers, I guess. I think they documented it somewhere). So higher average quality than something trained mostly on internet slop (twitter, reddit, etc). The only reason a normal LLM trained also on modern data is outputting less garbage is that they put a lot of effort into filtering the output. It is not difficult to find unfiltered local LLMs you can run on your own computer that will happily generate any kind of garbage text you want to (not limited to pre-1931 garbage!).

The first vintage language models I read about was end of last year. They trained models with cut-off dates at 1913 , 1929 , 1933 , 1939 , and 1946 (https://github.com/DGoettlich/history-llms/tree/main). Sounded great for wargame purposes. Unfortunately, as far as I can tell, they still have not released anything to the public, and might never, blaming issues with how to filter the output.

Sorry - only verified members can post on the forums.