• A
  • A
  • A
  • ABC
  • ABC
  • ABC
  • А
  • А
  • А
  • А
  • А
Regular version of the site

Predicting Grammatical Properties of Words Helps Us Read Faster

Predicting Grammatical Properties of Words Helps Us Read Faster

© iStock

Psycholinguists from the HSE Centre for Language and Brain found that when reading, people are not only able to predict specific words, but also words’ grammatical properties, which helps them to read faster. Researchers have also discovered that predictability of words and grammatical features can be successfully modelled with the use of neural networks. The study was published in the journal PLOS ONE.

The ability to predict the next word in another person’s speech or in reading has been described by many psycho- and neurolinguistic studies over the last 40 years. It is assumed that this ability allows us to process the information faster. Some recent publications on the English language have demonstrated evidence that while reading, people can not only predict specific words, but also their properties (e.g., the part of speech or the semantic group). Such partial prediction also helps us to read faster.

In order to access predictability of a certain word in a context, researchers usually use cloze tasks, such as 'The cause of the accident was a mobile phone, which distracted the ______'. In this phrase, different nouns are possible, but driver is the most probable, which is also the real ending of the sentence. The probability of the word 'driver' in the context is calculated as the number of people who correctly guessed this word over the total number of people who completed the task.

Student doing a cloze task on a smartboard
© Wikimedia Commons

The other approach for predicting word probability in context is the use of language models that offer word probabilities relying on a big corpus of texts. However, there are virtually no studies that would compare the probabilities received from the cloze task to those from the language model. Additionally, no one has tried to model the understudied grammatical predictability of words. The authors of the paper decided to learn whether native Russian speakers would predict grammatical properties of words and whether the language model probabilities could become a reliable substitution to probabilities from cloze tasks.

The researchers analysed responses of 605 native Russian speakers in the cloze task in 144 sentences and found out that people can precisely predict the specific word in about 18% of cases. Precision of prediction of parts of speech and morphological features of words (gender, number and case of nouns; tense, number, person and gender of verbs) varied from 63% to 78%.  They discovered that the neural network model, which was trained on the Russian National Corpus, predicts specific words and grammatical properties with precision that is comparable to people’s answers in the experiment. An important observation was that the neural network predicts low-probability words better than humans and predicts high-probability words worse than humans.

The second step in the study was to determine how experimental and corpus-based probabilities impact reading speed. To look into this, the researchers analysed data on eye movement in 96 people who were reading the same 144 sentences. The results showed that first, the higher the probability of guessing the part of speech, gender and number of nouns, as well as the tense of verbs, the faster the person read words with these features.

The researchers say that this proves that for languages with rich morphology, such as Russian, prediction is largely related to guessing words’ grammatical properties.

Second, probabilities of grammatical features obtained from the neural network model explained reading speed as correctly as experimental probabilities. ‘This means that for further studies, we will be able to use corpus-based probabilities from the language model without conducting new cloze task-based experiments,’ commented Anastasiya Lopukhina, author of the paper and Research Fellow at the HSE Centre for Language and Brain.

Third, the probabilities of specific words received from the language model explained reading speed in a different way as compared to experiment-based probabilities. The authors assume that such a result may be related to different sources for corpus-based and experimental probabilities: corpus-based methods are better for low-probability words, and experimental ones are better for high-probability ones.

Anastasiya Lopukhina, Research Fellow at the HSE Centre for Language and Brain.

Anastasiya Lopukhina, Research Fellow at the HSE Centre for Language and Brain.

Two things have been important for us in this work. First, we found out that reading native speakers of languages with rich morphology actively involve grammatical predicting. Second, our colleagues, linguists and psychologists who study prediction got an opportunity to assess word probability with the use of language model. This will allow them to simplify the research process considerably.

See also:

‘Speech, Facial Expressions, and Gestures Cannot Lie’

Would you like to know whether a speaker’s trembling voice or an accidental gesture can give them away? At HSE University in Nizhny Novgorod, researchers are developing an algorithm that analyses speech, facial expressions, and gestures, and determines whether information is truthful with 92% accuracy. The project has applications ranging from forensic examination and bank recruitment to fundamental research. Anna Khomenko, head of the research group and Senior Research Fellow at the Centre for Language and Brain at the HSE Faculty of Humanities in Nizhny Novgorod, explains how students and researchers are working together to create a corpus of video recordings, train a classifier, and prepare to introduce computer vision technology.

'We Would Like Our Corpora to Be Used More Widely'

The Linguistic Convergence Laboratory and the School of Linguistics at HSE University have created corpora of the Abkhaz-Adyghe languages spoken in the Western Caucasus. The corpora serve as valuable resources for studying these languages with their unique features and demonstrate the potential for their modern use. The corpora were developed through a series of field expeditions to the Caucasus conducted by HSE University researchers and students, combined with modern linguistic processing methods and collaboration with colleagues from regional universities. In this interview with the HSE News Service, Yury Lander, Leading Research Fellow at the Linguistic Convergence Laboratory and Associate Professor at the School of Linguistics, discusses the work of linguists.

‘In Science, You Are Your Own Boss’

Polina Nasledskova is interested in identifying gaps in linguistics and topics that have been overlooked by other researchers. In an interview for the  Young Scientists of HSE University project, she spoke about rare ordinal numerals in Nakh-Daghestanian languages, the benefits of knitting for concentration, and the beauty of the Patriarshy Bridge.

‘What Matters Is Not What You Study, but Who You Study with’

Katerina Koloskova began studying Arabic expecting to give it up after a year—now she cannot imagine her life without it. In an interview for the Young Scientists of HSE University project, she spoke about two translated books, an expedition to Socotra, and her love for Bethlehem.

Participants of HSE LED Conference Discuss Progress in Linguistics and Pedagogy

On April 20–21, the HSE School of Foreign Languages held the V International Scientific and Practical Conference ‘Languages. Education. Development’ (HSE LED). It was organised in an online format and dedicated to current trends in the development of modern knowledge in linguistics and pedagogy. Over two days, about 1,700 participants (including more than 220 speakers) took part in the event— 40% more than in the previous academic year.

'I Dream of Becoming Part of the International Semantics Community'

As a student, Stepan Mikhailov took part in an expedition to the Urals and became so deeply engaged that he eventually wrote his dissertation on a related topic—possessive constructions in the Khanty language. In this interview for the HSE Young Scientists project, he talks about bridging syntax and semantics, the importance of making time to cook and eat breakfast in the morning, and his favourite place in the village of Kazym.

HSE University Develops Tool for Assessing Text Complexity in Low-Resource Languages

Researchers at the HSE Centre for Language and Brain have developed a tool for assessing text complexity in low-resource languages. The first version supports several of Russia’s minority languages, including Adyghe, Bashkir, Buryat, Tatar, Ossetian, and Udmurt. This is the first tool of its kind designed specifically for these languages, taking into account their unique morphological and lexical features.

For the First Time, Linguists Describe the History of Russian Sign Language Interpreter Training

A team of researchers from Russia and the United Kingdom has, for the first time, provided a detailed account of the emergence and evolution of the Russian Sign Language (RSL) interpreter training system. This large-scale study spans from the 19th century to the present day, revealing both the achievements and challenges faced by the professional community. Results have been published in The Routledge Handbook of Sign Language Translation and Interpreting.

Mistakes That Explain Everything: Scientists Discuss the Future of Psycholinguistics

Today, global linguistics is undergoing a ‘multilingual revolution.’ The era of English-language dominance in the cognitive sciences is drawing to a close as researchers increasingly turn their attention to the diversity of world languages. Moreover, multilingualism is shifting from an exotic phenomenon to the norm—a change that is transforming our understanding of human cognitive abilities. The future of experimental linguistics was the focus of a recent discussion at HSE University.

Twenty vs Ten: HSE Researcher Examines Origins of Numeral System in Lezgic Languages

It is commonly believed that the Lezgic languages spoken in Dagestan and Azerbaijan originally used a vigesimal numeral system, with the decimal system emerging later. However, a recent analysis of numerals in various dialects, conducted by linguist Maksim Melenchenko from HSE University, suggests that the opposite may be true: the decimal system was used originally, with the vigesimal system developing later. The study has been published in Folia Linguistica.