HSE Scientists Develop Method to Compress Large Language Models Without Losing Quality

Researchers from the AI and Digital Science Institute at the HSE Faculty of Computer Science have developed a new compression method for large language models such as GPT and LLaMA that reduces their size by 25–36% without additional training or significant loss of accuracy. This is the first approach to use mathematical transformations—specifically, rotations of model weights—to make models more amenable to compression with structured matrices. The study results have been published in ACL Findings 2025. The code is available on GitHub.
Large language models such as ChatGPT and LLaMA demonstrate impressive results in text generation, translation, and other tasks, but their enormous size makes them costly to deploy and store. Traditional compression methods—such as reducing numerical precision, pruning redundant connections, or simplifying the architecture—often require time-consuming retraining and can degrade performance. Scientists sought a way to shrink model size quickly without compromising its intelligence.
Researchers from the Laboratory for Matrix and Tensor Methods in Machine Learning at the HSE FCS AI and Digital Science Institute proposed a method called ProcrustesGPT, based on the idea that a model’s output remains unchanged if special orthogonal transformations are applied to its internal weights—a kind of mathematical rotation. As the scientists explain, this is a transformation of space that can rotate or flip an image in any way but cannot stretch or compress any object. For example, if you take a piece of paper with a triangle drawn on it, you can flip or rotate it at any angle—the side lengths and the angles between them will remain exactly the same. In mathematics, such a transformation is called orthogonal. These transformations are chosen so that the model’s weights can be compressed more efficiently using structured matrices—mathematical constructions that require far less memory.
Ekaterina Grishina
Ekaterina Grishina, Research Assistant at the Laboratory for Matrix and Tensor Methods in Machine Learning, explains, 'Our work is based on an elegant mathematical concept—the Procrustes problem. Like the mythical figure Procrustes, who forced travellers to fit his bed, this method helps identify the optimal orthogonal transformation that reshapes the model’s weights into a simpler structure without distorting their essence. This idea inspired the name of our method, ProcrustesGPT, and became the key to achieving compression without significant loss of quality.'
As part of the study, the researchers tested two types of such structures: sums of Kronecker products and GS matrices. The method does not require additional training, works quickly, and can be applied to existing models. The experiments were conducted on the open OPT and LLaMA 2 models.
The new ProcrustesGPT method has demonstrated strong effectiveness: it reduces the size of large language models by about a third—more precisely, by 25–36% of their original size—while preserving their capabilities. The compressed models deliver results close to those of the originals, retaining 90–95% of their initial performance in generating coherent text and solving logical tasks.
Compared with other modern compression methods, such as SliceGPT—which also does not require lengthy additional training—ProcrustesGPT proved more accurate in most tests. This advantage is particularly evident for models in the LLaMA 2 family, where the proposed approach outperforms its counterpart by 9–10%.
Maxim Rakhuba
According to Maxim Rakhuba, Head of the Laboratory for Matrix and Tensor Methods in Machine Learning at the HSE AI and Digital Science Institute, 'Compression methods help accelerate the deployment of large language models on resource-constrained devices, such as mobile devices and IoT gadgets, making AI more accessible and widely integrated into everyday life.'
See also:
Biologists Discover Unique Properties of MiR-93-5p MicroRNA in Prostate Cancer
Researchers at the International Laboratory of Microphysiological Systems of the HSE Faculty of Biology and Biotechnology investigated how different isoforms of the same microRNA influence gene function in prostate adenocarcinoma. The study found that in some cases, microRNAs can reinforce each other’s effects by targeting and suppressing the same genes. This finding offers a fresh perspective on the molecular mechanisms underlying tumour development and on the search for disease biomarkers. The results have been published in PeerJ.
HSE Economists Use Search Queries to Forecast Birth Rates
Researchers from the HSE Faculty of Economic Sciences have shown that the accuracy of birth rate forecasts for Russia can be improved by almost 50% by incorporating the dynamics of online search queries related to pregnancy and childbirth into forecasting models. In the best-performing models, the forecasting error fell from 4.6% to 3.2%. The findings have been published in Populations and Economics.
When Looking at Their Own Faces, Men Forget Everything
In an experiment involving 15 healthy men, scientists at HSE University investigated how different phases of the cardiac cycle influence the excitability of the motor cortex when participants viewed either their own photograph or the faces of strangers. The researchers found that when participants looked at their own image, the brain’s response to signals from the heart was weaker, meaning that the influence of cardiac activity on the motor cortex decreased. This finding came contrary to expectations, as self-focused attention was thought to enhance the brain's sensitivity to internal bodily signals. The study has been published in Frontiers in Signal Processing.
HSE Researchers Discover Who Eats Out in Russia—And Why
Around one-third of Russians (31.3%) rarely eat out or buy ready-made meals. The core group of active consumers—those who eat out or purchase prepared food almost every day or several times a week—accounts for only about 9% of the population. These are the findings of a study conducted by the HSE Institute for Social Policy. According to the researchers eating out is no longer a marker of high social status in Russia.
Scientists Model How Interactions Between Societies Can Trigger Chaotic Behaviour
Scientists at HSE MIEM have proposed a mathematical model explaining how interactions between societies can influence their stability. Based on the classical theory of evolutionary games, the study reveals an unexpected effect: even a weak informational influence of one society on another can cause one society to remain stable while the other exhibits chaotic behaviour among its individual members. The study has been published in the International Journal of Bifurcation and Chaos.
Ancient Craniiform Brachiopod: A Newly Discovered Species with a Unique Shell Shape and Lifestyle
Scientists from HSE University, MSU, and Tallinn University of Technology have studied a fossil species of ancient brachiopods that lived in a warm sea in what is now northern Estonia more than 445 million years ago. These ancient brachiopods developed a cup-shaped shell with a protective 'cap' that shielded them from overgrowth by other marine organisms. The study has been published in Palaeogeography, Palaeoclimatology, Palaeoecology.
Scientists Develop Bacterium-Sized Microlaser
An international team of researchers, including scientists from HSE University–St Petersburg, has developed microlasers that emit deep-ultraviolet light at a wavelength of 255 nanometres. The devices operate at room temperature, and the smallest of them measures just two micrometres in diameter—roughly the size of a bacterium. These microlasers could be used in sensors, spectroscopic systems, photonic chips, and communication devices. The paper has been published in Optics & Laser Technology.
HSE Develops App for Assessing Phonological Processing in Children
Researchers at the HSE Centre for Language and Brain have developed a new digital tool for assessing children's phonological processing skills—the ZARYA (Sound Analysis of the Russian Language) test battery. It is the first standardised application in Russia designed to provide a fast and reliable assessment of children's ability to distinguish speech sounds, retain them in working memory, and perform phonemic analysis. The app runs on Android tablets and smartphones and is available for download from RuStore. Details of the test validation have been published in the Journal of Speech, Language, and Hearing Research.
Researchers Discover How Spelling Errors Slow Down Reading in Russian
Psycholinguists from the Centre for Language and Brain at HSE University–St Petersburg have shown that words that are frequently misspelled are processed more slowly by readers, even when presented with the correct spelling. The researchers confirmed this effect for the first time using Russian-language materials and found that response speed is most strongly linked to how confidently individuals can distinguish the correct spelling of a word from an incorrect one. The study has been published in The Mental Lexicon.
Scientists Discover Why Europium 'Misbehaves'
Europium is a rare-earth metal responsible for the pure red glow in displays and other luminescent materials. For a long time, however, it refused to emit light when surrounded by certain organic molecules known as acylpyrazolone ligands. Chemists have now uncovered the reason: in europium complexes with these ligands, a 'black window' appears—a charge-transfer state in which the energy absorbed by the ligand is dissipated as heat rather than emitted as light. Understanding this mechanism opens the way to designing more efficient red-emitting materials for displays, fluorescent thermometers, and chemical sensors. The results have been published in Dalton Transactions.


