AI and the Myth of the Total Archive

History Beyond the Web

Over the summer, Microsoft published some research suggesting that historians were amongst the top jobs most at risk from being done by AI.

It was similarly bad news for mathematicians, data scientists, translators, and radio DJs. Things looked better for phlebotomists, roofers, and embalmers, who were close to the bottom of the list.

Rank Job Title Coverage Score
1 Translators 0.98 0.49
2 Historians 0.91 0.48
3 Passenger Attendants 0.80 0.47
4 Sales Representatives 0.84 0.46
5 Writers and Authors 0.85 0.45
6 Customer Service Representatives 0.72 0.44
7 CNS Tool Programmers 0.90 0.44
8 Telephone Operators 0.80 0.42
9 Ticket Agents 0.71 0.41
10 Broadcast Announcements and Radio DJs 0.74 0.41
Top 10 occupations with highest AI applicability score. “Working with AI”, arXiv, 2025.

No one likes to be told they’re replaceable, especially by a machine. Naturally, I was a bit perturbed.

But this got me thinking—what exactly does Microsoft think historians do?

To be fair to the researchers, they used a standardised list of job activities produced by the Occupational Information Network (O*NET). Nonetheless, the results are pretty revealing in terms of how technology companies think about work itself.

Virtual Historians

I accept that my job as a historian requires less physical labour and skill than, say, a phlebotomist or a roofer. So in some sense, Microsoft is right. My work, like a lot of jobs today, involves sitting in front of a screen for long periods of time. I read and write, but most of the time, I do those things on a computer. (Although I still take all my notes by hand.)

But there’s a fiction at play here. Historians don’t actually live in some disembodied virtual reality. Perhaps more than some other professions, we actually spend a lot of time doing real things in the physical world.

Teaching is obviously a big part of that. Giving a lecture or discussing something in a seminar requires a variety of physical skills that, as anyone who tried to teach online during COVID will tell you, are not easily replicable in a digital environment. That’s especially the case if you’re making use of objects, or art, or even re-creating historical experiments in your teaching.

Now, I concede that teaching could be made digital. Yes, it might not be as interesting or engaging, but seminars and lectures can be taught online. And in fact, that’s part of the trick of technology—to make you settle for less.

Still, there’s another obvious part of historical research which cannot be done by AI. It’s called “going to an archive”. Yes, historians get in a car, or on a bus, or sometimes on a plane, and go to a library or alike. (I’m actually lucky enough that I can walk to the Modern Records Centre at the University of Warwick, just across from my office.) Once we get there, we turn the pages of old bits of paper, and begin reading, transcribing, and taking notes. AI can’t do that for one simple reason—this stuff isn’t on the Internet.

Archives of the World Federation of Scientific Workers, Modern Records Centre, University of Warwick.

Archives of the World Federation of Scientific Workers, Modern Records Centre, University of Warwick.

The Myth of the Total Archive

This gets to a bigger problem with AI which I refer to as “the myth of the total archive”. As we all know, large-language models (LLMs) are trained on essentially everything available on the Internet. That’s a lot of information. Estimates vary, but there might be around 175 zettabytes out there. That’s 175 trillion GB.

But that doesn’t cover all of human knowledge, past and present. Not even close. If we just look at the English language, there must be billions words of text that are not on the Internet. I listened to a brilliant talk yesterday where my colleague showed a photograph of the heap of documents she had discovered in the back of a police station in Uganda. That’s definitely not online. Similarly, the majority of sources I used for my PhD, including thousands of nineteenth-century letters, are not digitised.

Then, if we start looking beyond English, the world of the non-digital gets even bigger. Thousands of palm-leaf manuscripts, in languages such as Sanskrit and Kannada, are piled up in temples across India. These, like the vast majority of the world’s manuscripts, have not been digitised and, frankly, never will.

You can add to this the problem of censorship. If you think the Chinese Communist Party is going to digitise all the records of its activities from 1949 to 1989 and put them on the Internet, then think again. And it’s not just the communists. The British state spent decades trying to keep the records of its activities during the Mau Mau rebellion secret, hoarding them at Hanslope Park. Same goes for more and more countries around the world these days. The records of the state are increasingly kept behind closed doors.

On top of that, historians don’t just deal with texts. Over the past few decades, we’ve made increasing use of images and objects—what we call “material culture”. I spent a good deal of my PhD looking at the skull collection held in the Edinburgh University Anatomy Department. Others study old scientific instruments, or cotton textiles, or even the entire architecture of a temple complex. None of this is easy to accurately digitise. And even if you did, the volume of stuff out there is vast.

If I had to guess, I’d say that LLMs have likely been trained on less than half of the world’s entire textual output. And if we expand that to include historical sources beyond the textual, LLMs have barely scratched the surface.

The Internet Isn’t Everything

And so this is the myth of the total archive. It’s the myth that AI has been trained on everything. When in fact, AI has experienced but a fraction of human culture. It’s the myth that allows AI researchers to suggest that historians’ work could done by AI.

It’s also a myth that tech companies like Microsoft have an interest in promoting. The more we believe that the world is contained on the Internet, the less reason we have to log off and leave our desks. That’s good for Microsoft’s bottom line, but not for the rest of us.

So maybe it’s time to close Word, and get out there.