HSE Researchers Create New Corpus of Early Child Speech in Russian

Researchers at the HSE Centre for Language and Brain have presented RusLan-M, an open multimedia corpus that makes it possible to trace the development of early child speech in Russian from first words to the emergence of complex grammatical constructions. The database contains around 41 hours of video recordings and more than 35,000 child utterances. The new resource will help researchers study more precisely how children acquire Russian and, in the longer term, develop more reliable tools for assessing speech development. The study has been published in Language Resources and Evaluation.
Despite the widespread use of Russian, data on its early acquisition remain limited. The international CHILDES database brings together hundreds of collections covering more than 40 languages, but Slavic languages are significantly underrepresented. At the time the study was prepared, only two corpora of children’s Russian speech were available in CHILDES. The new RusLan-M resource is intended to address this gap by providing researchers with detailed, standardised, and accessible data on how children acquire Russian.
The corpus is based on observations of two Russian-speaking children, Tosya and Yasha. Researchers at the HSE Centre for Language and Brain recorded the development of their speech over an extended period, rather than comparing children of different ages at a single point in time. The collection comprises 288 recordings: Tosya was recorded from the age of 10 months to 3 years and 10 months, while Yasha was recorded from 1 year and 4 months to 3 years of age. The total duration of the video recordings was 2,454 minutes, or around 41 hours, while the number of child utterances reached 35,386.
All recordings were accompanied by transcripts in the international CHAT format and published on TalkBank, an international resource for storing and analysing speech data. When publishing the corpus, the researchers complied with requirements concerning personal data protection and the anonymisation of personal information.
A key feature of the new resource is that it makes it possible to study children’s speech in a natural environment. Unlike laboratory experiments or parental questionnaires, the video recordings capture children’s spontaneous interactions with people close to them and their everyday linguistic environment. This makes it possible to analyse not only what children say, but also the speech context in which their language develops.
Mariia Diachkova
‘It was important for us to create a resource that would allow us to see the development of children’s speech not as a set of individual test results, but as a living and continuous process. Thanks to the longitudinal recordings, we can trace how an individual child’s speech gradually changes, which grammatical constructions emerge earlier and which later, and how the linguistic environment influences development. Such open and comprehensively annotated data is particularly important for Russian,’ said Mariia Diachkova, Research Assistant at the HSE Centre for Language and Brain and one of the study’s authors.
Pilot studies have demonstrated how RusLan-M data can be used to analyse the development of morphology and syntax over time. The corpus makes it possible to track the emergence and increasing complexity of specific linguistic constructions, compare different stages of speech development, and formulate new hypotheses about the mechanisms of language acquisition.
Valeriia Lelik
According to the authors, the significance of such data extends beyond fundamental linguistics. ‘Corpus-based measures can be used to develop normative benchmarks for speech development at different ages. In the future, this could help specialists distinguish more accurately between individual patterns of language acquisition and potential difficulties, as assessments would be based on data from children’s real spontaneous speech,’ notes Valeriia Lelik, Research Assistant at the Centre for Language and Brain and one of the study’s authors.
The authors emphasise that the new corpus partly addresses the significant shortage of open data on the development of the Russian language.
Svetlana Dorofeeva
‘Diary records of child speech have been used as research material in Russian developmental linguistics for decades. Everyone working in this field knows how much effort is required to create corpora of child speech. However, it is important to note that no multimodal corpora of child speech in Russian comparable to RusLan-M—incorporating audio, video, and text and featuring detailed linguistic annotation—had previously been publicly available. Thanks to this large-scale effort, the research community now has access to this valuable tool,’ said Svetlana Dorofeeva, project lead and Senior Research Fellow at the HSE Centre for Language and Brain.
The corpus was created through the joint efforts of a large team of parents, linguists (both researchers and students), and programmers. It brings together first-hand parental experience, academic expertise, and advanced digital technologies.
Olga Dragoy
‘In the longer term, expanding collections of this kind will make it possible to compare the speech development of a larger number of children, identify age-related patterns, and create more accurate models of language acquisition,’ said Olga Dragoy, one of the study’s authors and Director of the HSE Centre for Language and Brain.
RusLan-M also demonstrates how corpus linguistics methods can transform years of observations of children’s speech into a resource for a wide range of research, from theoretical linguistics to psychology and the practical assessment of speech development.
See also:
HSE University to Develop Predictive Analytics System for Icebreaker Motors
Industrial automation is one of the key applications of artificial intelligence. A predictive analytics system for large electric motors is among the solutions being developed for the industry as part of HSE University’s Strategic Technological Project ‘Multi-Agent Platform of AI Solutions for Industry-Specific Tasks.’ What is predictive analytics, how can it improve the operation of electric motors, and what specialists joined forces to develop this technology? Anton Zarubin, Dean of the School of Computer Science, Physics, and Technology at HSE University–St Petersburg and the project development coordinator, explains in this interview with the HSE News Service.
How to Assess Students’ Knowledge in the Age of AI
A researcher at HSE University has proposed a flowchart to help lecturers decide how to assess students who use artificial intelligence. It shows where the use of AI should be restricted and where it can be incorporated into the learning process. The article has been published in IT Professional.
Scientists Train Neural Network to Generate Process Plans from 3D Models
Researchers at the HSE FCS AI and Digital Science Institute have developed CAD2TechSpec, a framework that converts 3D models of mechanical parts into machining process plans—step-by-step instructions for machine tools. The solution aims to reduce the time required for the design and preparation of technical process documentation in mechanical engineering, aircraft manufacturing, and other high-tech industries. The study findings have been published in PeerJ Computer Science.
Biologists Discover 'Molecular Fingerprint' of Preeclampsia
Researchers at HSE University employed a new method to model hypoxia in placental cells during pregnancies complicated by preeclampsia and identified molecular markers of tissue hypoxia. Since hypoxia is one of the key mechanisms underlying preeclampsia, these findings are important for a more accurate and timely diagnosis of the disease and for the development of effective treatment methods. The paper has been published in Placenta.
‘Hedgehog’ Versus ‘Relatives’: Researchers Measure How the Brain Responds to Unexpected Words During Natural Speech
Russian neurophysiologists, including researchers from HSE University, have demonstrated the feasibility of using event-related fields (ERFs) to study brain activity during natural speech perception. The researchers showed that this approach can be applied not only to individual words but also to continuous speech. Their findings indicate that words whose meanings differ significantly from the preceding context require longer processing times. The study also reveals that the brain processes function words in two stages: first, it identifies their grammatical role and then uses this information to predict the next word. The study has been published in Frontiers in Human Neuroscience.
Scientists Develop Algorithm for More Reliable Processors in Data Centres
Researchers from HSE MIEM and Samara University have developed the LRF-3D algorithm to automatically bypass idle nodes in three-dimensional networks-on-chip. Thanks to its hierarchical architecture, the algorithm outperforms existing solutions in both speed and path accuracy, improving processor reliability for use in data centres, supercomputers, and AI computing. The source code and test results are publicly available.
Researchers Rank Recommendation Algorithms Using Sports Tournament Model
Researchers from the AI and Digital Science Institute at the HSE Faculty of Computer Science have developed an approach for selecting recommendation algorithms more effectively. Their approach uses pairwise comparisons of algorithms to create a tournament table, with the overall ranking based on their performance across all datasets in the tournament. This can reduce the number of algorithms that need to be tested when developing new services, saving both time and money. The study was presented at the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 2026).
Researchers Develop Method for Direct Generation of Regulatory DNA
Researchers at HSE University have developed a model for generating promoters and enhancers—DNA sequences that regulate gene activity. The model works directly with DNA nucleotides, without first transforming them into a continuous numerical representation. This solution could be useful for applications in synthetic biology and gene therapy. The study results were presented at the ICLR 2026 Workshop ‘Generative AI in Genomics (Gen^2): Barriers and Frontiers.’
Researchers at HSE University and Sber Train Neural Networks to Better Predict User Preferences
The HSE FCS AI and Digital Science Institute and Sber have introduced a new architecture for recommendation systems that combines two classes of models, enabling algorithms to better predict users’ interests and needs. A preprint of the paper has been published on arxiv.org and presented at Urban ML.
Physicists Discover What Happens Inside a Stable Vortex
Large vortices with characteristic spiral arms are often observed in the atmosphere and the ocean. Physicists from HSE University have explained how these structures form and why they retain their shape. The researchers found that velocities at points located along the same vortex arc remain correlated even over long distances. At the same time, this correlation weakens rapidly with increasing distance from the vortex centre. These differences help explain the formation of spiral arms and may improve models of atmospheric and oceanic currents. The findings have been published in Physical Review Fluids.


