• A
  • A
  • A
  • ABC
  • ABC
  • ABC
  • А
  • А
  • А
  • А
  • А
Regular version of the site
  • HSE University
  • News
  • Teaching a Machine to Read the Past: HSE Develops Neural Network to Decipher Manuscripts

Teaching a Machine to Read the Past: HSE Develops Neural Network to Decipher Manuscripts

Manuscript of playwright Aleksandr Sukhovo-Kobylin

Manuscript of playwright Aleksandr Sukhovo-Kobylin
Source: Russian State Archive of Literature and Art

Diaries and letters are an invaluable resource for humanities scholars. But what can be done when the text is impossible to read? At the HSE Faculty of Humanities, this challenge has been translated into the language of mathematics: a team of philologists, historians, and machine learning specialists has created an information system that not only recognises illegible handwriting but also helps analyse archival content.

The Background

Work with handwritten sources has long been a tradition at the faculty. A new technological stage of this began in 2019, when HSE joined the Autograph project under the leadership of Elena Penskaja (now Head of the Centre for Digital Archival Studies at the Faculty of Humanities). The project itself was launched in 2014 by a group of researchers from the Russian State Archive of Literature and Art. Almost immediately, the initiative—which enabled students, scholars, and literature enthusiasts worldwide to study digital copies of manuscripts—received support from the Pushkin House and the Russian Science Foundation.

In 2022, a group of Autograph participants decided to take the work further. They submitted a new application to the Russian Science Foundation and won a grant for the interdisciplinary, inter-university project ‘Russia’s Cultural Heritage: Intelligent Analysis and Thematic Modelling of a Corpus of Handwritten Texts.’ Historians, mathematicians, and philologists from HSE joined forces with their long-standing partners from Tomsk State University.

Elena Penskaja
Photo: HSE University

The goal was ambitious: to develop digital tools capable of transforming chaotic collections of manuscripts—diaries, letters, and ego-documents from the nineteenth and early twentieth centuries—into structured data using machine learning algorithms. The task was not merely digitisation, but the automatic identification of hidden themes, narratives, and meanings, alongside the cataloguing and intelligent analysis of archival materials.

The project formally concluded in 2025, but the research continues. Elena Penskaja and Candidate of Physical and Mathematical Sciences Nikita Lomov have created a functioning information system whose primary mission is to teach machines to read the unreadable.

How It Works: Lines, Entities, and the YOLO-HTR Neural Network

Traditional manuscript cataloguing in archives and libraries is based on organising documents into fonds, sheets, storage units, and page numbering. Digitisation adds image-based navigation, which is useful but does not solve the core problem: the text itself remains unrecognised.

The system developed at the Faculty of Humanities goes two steps further. It employs the original YOLO-HTR (You Only Look Once + Handwritten Text Recognition) neural network architecture, which simultaneously performs two tasks: locating lines of text on an image and deciphering them. As a result, each line of a manuscript becomes linked not only to a page number, but also to its textual content.

But this is only half the task. The key innovation is semantic navigation. Using large language models, the system identifies so-called entities within the text: not only traditional categories such as ‘persons,’ ‘locations,’ or ‘organisations,’ but also more complex ones such as ‘state of health,’ ‘political event,’ or ‘reflection.’ A user can click on any entity and instantly access all lines and pages where it is mentioned. This transforms an archive from a mere stack of images into an interconnected knowledge base with bidirectional cross-references.

‘We achieve content-based archive organisation,’ explained Nikita Lomov. ‘From subjects of interest, users can navigate directly to specific lines and pages, while each page and its lines provide a line-by-line list of referenced entities.’

Sukhovo-Kobylin’s Diaries: A Challenge That Lasted 40 Years

One of the most striking case studies involves the diaries of playwright Aleksandr Sukhovo-Kobylin (1817–1903). A mysterious figure, he was once suspected of murdering his French lover, wrote three plays that entered the Russian literary canon, and published almost none of his diaries during his lifetime.

Despite their impressive volume, the diaries themselves have only been partially published. Deciphering the published portion took around 40 years, and even that edition contains omissions and inaccuracies. Sukhovo-Kobylin’s handwriting is so illegible that it can easily confound an untrained reader.

Manuscript of playwright Aleksandr Sukhovo-Kobylin
Source: Russian State Archive of Literature and Art

The Faculty of Humanities team uploaded 380 pages of diaries into the system—more than 10,000 lines of text, around 5,000 of which had existing published transcriptions used to train the neural network. By comparison, the system recognises the handwriting of Fyodor Litke (Friedrich Benjamin von Lütke) and Modest Korf with an error rate of just 3–5% at the letter level. For Aleksandr Sukhovo-Kobylin, the error rate rises to 10% for letters and 28% for words.

Even so, the developers emphasise that this result represents a major breakthrough for researchers. Most errors can be corrected easily, while the text becomes highly readable in places where scholars previously had to spend several minutes deciphering each word.

Dialogue with the Machine: How to Question an Archive

Modern large language models (such as ChatGPT, DeepSeek, Gemini, and Claude) already allow users to interact with them almost as though they were speaking to a human interlocutor. The Faculty of Humanities developers have gone a step further by adapting this format specifically for archival work.

Researchers can formulate queries in natural language—for example, ‘show all mentions of illnesses in the 1850s’ or ‘identify journeys with routes and companions’—and the system will return not a continuous text, but a structured dataset with fields suitable for further analysis. This makes it possible to track the temporal dynamics of references, identify co-occurrences of entities, and reconstruct social networks, conflicts, and patterns of movement.

‘Much depends here on the wishes and ambitions of our colleagues in philology,’ said Nikita Lomov. ‘It is precisely their research interests that will determine new types of supported queries and drive the expansion of our system’s capabilities.’

Scaling Up and the Scholarly Community

The developers envision two parallel paths for the system’s future development.

The extensive path involves expanding data volumes by creating similar systems for other collections of ego-documents, particularly those for which textual transcriptions already exist.

The intensive path focuses on improving the algorithms themselves: reducing recognition errors, lowering the need for annotated data, and achieving more precise entity extraction even when transcriptions are imperfect.

However, the most important condition for success is the emergence of a community of engaged users. At present, the system operates primarily in a research mode. To seriously consider large-scale expansion, hundreds of active researchers, historians, philologists, and students are needed—not merely to observe, but to formulate queries, propose new categories of entities, and test hypotheses.

‘We would like to see a real community form around systems like this,’ said Nikita Lomov. ‘Only when the number of genuinely interested users reaches into the hundreds can the question of scaling be seriously addressed.’

The information system in Russian is already available online (currently demonstrated through the example of Sukhovo-Kobylin’s diaries).

The project continues under HSE University’s 2026 Fundamental Research Programme (‘Language, Literature, and Culture in Historical and Social Dimensions’). Those interested in testing the system or collaborating can join through the Centre for Digital Archival Studies at the HSE Faculty of Humanities.

See also:

‘Working with AI Solves a Wide Range of Engineering Problems’

Artificial intelligence is a working tool based on a balanced combination of algorithms and engineering. Experts and doctoral students from the HSE Moscow Institute of Electronics and Mathematics explain how AI technologies can improve an application, device, or system, and what engineering tasks are solved in the process.

‘The Peak of Stupidity’ and ‘The Valley of Despair’: HSE Economists Propose an Explanation for the Dunning–Kruger Effect

The Dunning–Kruger effect, which describes a sharp surge in self-confidence among beginners followed by an equally rapid decline as they gain experience, can be explained by the nature of the learning process and the acquisition of new knowledge. This conclusion was reached by Andrey Vorchik of the HSE Faculty of Economic Sciences together with independent researcher Murat Mamyshev. They developed a mathematical model of learning and demonstrated how subjective confidence is formed and changes as knowledge accumulates, as well as how teachers can reduce the ‘valley of despair’ experienced by learners.

Advancing Collaboration: HSE Faculty of Computer Science and Harbin Institute of Technology Hold Joint Seminar

From July 13 to 16, 2026, the Faculty of Computer Science hosted the Russian–Sino Research Seminar on Machine Learning Applications, organised by the HSE Laboratory for Cloud and Mobile Technologies in partnership with the Harbin Institute of Technology (China). A delegation comprising seven students and three university representatives travelled to Moscow to take part in an intensive four-day programme.

A New Section on AI and a Prizewinning Paper: Early-Career HSE Researchers Take Part in IEEE EDM Conference

The 27th IEEE International Conference of Young Professionals in Electron Devices and Materials (EDM) has taken place in the Altai Republic. This year, researchers from HSE University presented the results of their research and were involved in organising a new section on artificial intelligence. A paper by HSE master’s student Rodion Sidorenko was awarded third place in the research paper competition at the conference.

HSE Initiates Development of Ethical Standard for Anthropomorphic Robots

Beyond technological solutions, the development of anthropomorphic robotics also demands ethical ones. In July 2026, the HSE Institute for Robotics Systems hosted a foresight session dedicated to developing an Ethical Standard for Anthropomorphic Robotic Complexes. Representatives from businesses, government bodies, scientific organisations, and universities gathered to discuss key ethical and legal issues surrounding the development of anthropomorphic robotic complexes. The main outcome of the meeting was a draft of the Ethical Standard.

Physicists at HSE University and FIAN Discover Way to 'Photograph' Sound for Testing Materials Used in 6G Communications

Researchers at HSE University, in collaboration with colleagues from the Lebedev Physical Institute of the Russian Academy of Sciences (FIAN), have developed a method for rapidly determining how firmly a film is bonded to a substrate. This is important for the creation of ultrahigh-frequency acoustic filters, which are key components of next-generation 5G and 6G communications. For the first time, researchers have succeeded in measuring the lateral rigidity of the bond between a two-dimensional material film and a substrate in this way. The study results have been published in Applied Physics Letters.

HSE University to Launch New AI Supercomputer

HSE University is preparing to launch its second supercomputer. The new cluster will be primarily dedicated to artificial intelligence (AI) workloads and will complement the existing cHARISMa supercomputer. It is scheduled to become operational by the end of 2026.

HSE MIEM Students to Develop Two Satellites from Scratch for Orbital Experiments

The devices, created by student teams, will conduct space research on the properties of promising solar cells, on-board energy storage systems, and serial electronics for student satellites.

Biologists Discover Unique Properties of MiR-93-5p MicroRNA in Prostate Cancer

Researchers at the International Laboratory of Microphysiological Systems of the HSE Faculty of Biology and Biotechnology investigated how different isoforms of the same microRNA influence gene function in prostate adenocarcinoma. The study found that in some cases, microRNAs can reinforce each other’s effects by targeting and suppressing the same genes. This finding offers a fresh perspective on the molecular mechanisms underlying tumour development and on the search for disease biomarkers. The results have been published in PeerJ.

Researchers Discover How Spelling Errors Slow Down Reading in Russian

Psycholinguists from the Centre for Language and Brain at HSE University–St Petersburg have shown that words that are frequently misspelled are processed more slowly by readers, even when presented with the correct spelling. The researchers confirmed this effect for the first time using Russian-language materials and found that response speed is most strongly linked to how confidently individuals can distinguish the correct spelling of a word from an incorrect one. The study has been published in The Mental Lexicon.