1831 1930

Chronologic-EN: how well can language models represent the past?

Researchers would like to study the past. Ordinary people would like to converse with it. Can language models help in either case?

Ted Underwood, Ziliang Qiu, Sarah Griebel, Laura K. Nelson, Edwin Roland, Wenyi Shang, Matthew Wilkens · September 2026

Historians and social scientists face a fundamental constraint: history only happens once. We cannot rerun the French Revolution, or survey past populations to understand variations in their opinions. Language models excite researchers because they seem to open a new avenue of inquiry that gets around this constraint. If we had a model that represented the space of probability governing a period's writing (a big if!), historians could for the first time run reproducible experiments:

Meanwhile, people who aren't historians just like the idea of having a conversation with another era. What does a voice from the past sound like? How does it respond to new ideas?

Why a benchmark is needed

Scholars already know there are reasons to be cautious here. Our record of the past is biased and full of omissions. And reconstructing language is not the same thing as reconstructing the people who produced it — especially because language models have a tendency to homogenize different perspectives.

What we don't yet have is a way to measure success or failure. We may suspect that commercial language models are caricaturing the past or flattening variation, but our evidence is anecdotal. Flying blind also makes it difficult to improve performance. For instance, several recently-released models (e.g. Talkie-1930) were trained only on sources before a cutoff date in the early twentieth century. Does that limitation improve authenticity? What gets lost? Is the trade-off worthwhile? It's hard to say.

Why is there no benchmark for historical representation yet? While testing factual knowledge about the past might be a fairly simple task, testing a model's ability to respond like a writer from the past is harder. The benchmark we present here only covers one century (1831-1930), in one language (English), but took us a year to construct.

Why representation of the past is hard to measure

To test mathematical reasoning, or factual knowledge about history for that matter, we can gather living experts, get them to write questions with known correct answers, and then give models a multiple-choice test.

None of that is possible if we want to measure a model's ability to speak from a vantage point located in the past. There are no living experts we can poll to get correct answers: while historians know a lot about the past, they are not trained to respond like astronomers, fashion writers, and diarists of the 1870s. Instead, we have to extract ground truth answers from period documents. The differences between answers are often the interesting part, so we pair each question with a "metadata frame" that indicates the period and genre of the text, and often the social position of the writer, to specify whose attitudes are being represented.

But even with social context specified, many questions have multiple valid answers. A ground truth answer can tell us that one Unitarian writer believed this in 1885. But other Unitarians might have answered the question differently. How can we estimate the penumbra of acceptable variation around each ground truth example?

Answering that question took us more than a year. The solution we finally settled on was to locate more than one ground truth answer for 10% of our questions, and use the variance between ground truths to estimate the acceptable range of variation for other tricky questions. (We measure variation in answer quality in a space of Elo-like latent strengths determined by pairwise comparisons between answer options; see section 5.2 of the paper for a full description of the method.)

Along the way, we discovered that multiple choice format is largely useless here. Reasoning models turned out to be much better at recognizing correct answers than at generating them, which made multiple choice unrealistically easy. On the upside, this made automated judging more effective than we had initially expected. A moderately strong reasoning model can accurately judge the output of models stronger than itself.

Where models stand

The other important thing we found is that models have been getting better at representing the past. This is somewhat surprising, because we've had no benchmarks in this space, and it hasn't — as far as we know — been a major focus of commercial research. But we do see dramatic improvement. Scores on the part of the benchmark that asks models to generate text for a specific genre and social context, for instance, increase from 26% in August 2024 (GPT-4o) to 72% in August 2026 (GPT-5.6). See section 6 of the paper for other results. Full explanation of the improvement will take more research, but the data we see so far are consistent with a hypothesis that general improvements in reasoning also make models better at ventriloquizing the past.

Horizontal bar chart comparing six language models on the Chronologic-EN generation score, plotted against the full 0 to 100 range of the benchmark. Bars are ordered from lowest score at the bottom to highest at the top: Qwen 72B it, 20.2; GPT-4o, 26.0; Talkie 1930 13B it, 32.9; GPT-4.1, 44.0; Kimi K3, 66.6; and GPT-6 Astra, 71.4. The two most recent frontier models lead by a wide margin, roughly doubling the scores of the oldest models tested. Talkie 1930 13B it, the period-specific model trained only on text from before 1930, falls in the middle of the group, ahead of GPT-4o and Qwen 72B it but well behind GPT-4.1, Kimi K3 and GPT-6 Astra. Every model scores below 72, leaving more than a quarter of the scale unclaimed.
Figure 1. Free-text evaluation of answers to knowledge and constrained generation questions, along with the score that measures a model's stylistic fidelity to the target date provided in a question. These are just three of five scores Chronologic-EN can produce; we focus on them here because they dramatize development over the last two years, and the strengths of date-restricted models like Talkie 1930 13B relative to (likely) larger competitors such as GPT-4o. Closed reasoning models do lead on free-text evaluation overall, but there are questions a closed model cannot be used to answer.

While we can also compare commercial frontier models to period-specific models (like Talkie-1930), it might be better to say that these two classes of tools are attempting different things. Commercial models are definitely better at free-text generation, in part because they can understand complex questions and work through long chains of reasoning. But they cannot, in principle, answer some questions researchers want to pose, because they don't measure the likelihood of a given response in a given context. Among models capable of reporting a likelihood, Talkie-1930-base is the strongest.

What happens next

None of the models we tested are perfectly reliable. Commercial frontier models are not challenged by knowledge questions, and are surprisingly good at matching their responses to the prose style of a given decade. But they still struggle with questions that ask them to respond like a writer (or fictional character) in a specific social context.

Models trained only on period text struggle with a wider range of questions. In part this is because they haven't been trained to follow complex instructions as extensively as commercial models. But it's also true, as Benjamin Breen has pointed out, that removing texts after 1930 doesn't in itself give a model a specific historical perspective. Talkie-1930 floats across a long timeline, occupying a more diffuse vantage point than any living human being ever did.

Once language models can reliably represent specific points of view, it will make sense to ask whose points of view they represent, and who has been left out. But at the moment, models still confront a more fundamental challenge. Open-weight models aren't reliably representing anyone's perspective; they tend to produce a kind of averaged, context-free voice. Frontier reasoning models can do better, but still fail through excess self-consciousness: often overplaying a role in ways that telegraph a twenty-first-century understanding of its limitations.

We have designed this benchmark to measure models’ ability to inhabit a specific, limited point of view, because that’s one of the first challenges we need to address before models can be reliable tools for historical research. But it won’t be the last challenge! Historical representation is not a one-dimensional problem that can be summed up with a single measure of accuracy (when you reach 100%, problem solved.) On the contrary, we feel confident that our measure needs to be complemented by others — benchmarks that cover other periods and languages, but also, just as importantly, benchmarks that define the problem of representation in different ways. Chronologic-EN-1.0 is only an opening move in what we hope will be a long conversation.

References

Benjamin Breen, “Are ‘vintage LLMs’ the start of a new humanistic field?”, Res Obscura, April 2026.

Nick Levine, David Duvenaud, Alec Radford, “Introducing talkie: a 13B vintage language model from 1930,” April 2026.