• A
  • A
  • A
  • ABC
  • ABC
  • ABC
  • А
  • А
  • А
  • А
  • А
Regular version of the site

Transformer-Driven Approaches for Emotion Detection in Social Media Content

Student: Sanjar Javodov

Supervisor: Vasilii Gromov

Faculty: Faculty of Computer Science

Educational Programme: Data Science (Master)

Year of Graduation: 2026

The rapid growth of social media has made automatic analysis of emotional expression an important task in natural language processing. While sentiment analysis usually focuses on broad polarity categories, emotion detection aims to identify more fine-grained affective states such as joy, anger, sadness, fear, surprise, and disgust. This task is especially challenging in social media environments, where texts are short, informal, context-dependent, and often contain slang, irony, hashtags, emotionally loaded expressions, and domain-specific references. This thesis investigates transformer-driven approaches to emotion detection in social media content, focusing on football-related discourse as a domain-specific case of emotionally rich online communication. The research addresses a gap in existing studies that often rely on general-domain datasets, report only in-domain benchmark results, or give limited attention to corpus construction, class imbalance, annotation reliability, model interpretability, and cross-domain generalisation. Football-related social media discourse is selected because it combines large-scale public engagement with event-driven emotions, rivalry-based communication, fan reactions, and a difficult boundary between factual reporting and emotional expression. The empirical part of the thesis includes the construction and preprocessing of a football-related social media corpus, the development of a seven-class single-label emotion annotation scheme, and the comparative evaluation of classical and transformer-based classification models. The experiments include majority and TF-IDF baselines, Logistic Regression, Naive Bayes, Linear SVM, RoBERTa-based models, BERTweet, focal loss, class-weighted training, weighted sampling, and hyperparameter tuning. The best transformer configuration, RoBERTa-base with sqrt-inverse class-weighted cross-entropy, achieved the highest test macro-F1 score of 0.632, while a tuned TF-IDF + Linear SVM baseline remained highly competitive with a macro-F1 score of 0.599. McNemar’s exact test showed that the final RoBERTa model was not significantly different from the tuned SVM in paired accuracy, indicating that its main advantage lies in more balanced class-wise performance rather than in a large gain in overall correctness. The thesis extends the original evaluation in three additional directions. First, an inter- annotator agreement audit was conducted on a 300-post subset, producing overall agreement of 0.350 and Cohen’s Kappa of 0.212, which confirms the ambiguity of fine-grained emotion annotation in short sports-related posts. Second, explainability was expanded through both LIME-based local explanations for RoBERTa predictions and SHAP-style global lexical attribution for the tuned TF-IDF + Linear SVM baseline. Third, a cross-domain validation experiment tested whether models trained on football transfer to an independently labelled volleyball corpus. Both models showed substantial degradation under zero-shot transfer, but RoBERTa transferred better than SVM: RoBERTa macro-F1 decreased from 0.604 on the football test set to 0.373 on volleyball, while SVM macro-F1 decreased from 0.599 to 0.283. The thesis contributes to the field by presenting a domain-sensitive workflow for emotion detection in football-related social media, systematically comparing classical and transformer-based approaches under class imbalance, combining statistical testing with error analysis and explainability, and evaluating cross-domain generalisation between sports domains. The findings show that transformer-based models are effective for domain-specific emotion detection, but their performance depends strongly on annotation quality, imbalance handling, and the linguistic characteristics of the target domain. Overall, the thesis demonstrates that successful emotion detection in social media requires strong models, careful corpus design, transparent evaluation, and domain-aware interpretation.

Student Theses at HSE must be completed in accordance with the University Rules and regulations specified by each educational programme.

Summaries of all theses must be published and made freely available on the HSE website.

The full text of a thesis can be published in open access on the HSE website only if the authoring student (copyright holder) agrees, or, if the thesis was written by a team of students, if all the co-authors (copyright holders) agree. After a thesis is published on the HSE website, it obtains the status of an online publication.

Student theses are objects of copyright and their use is subject to limitations in accordance with the Russian Federation’s law on intellectual property.

In the event that a thesis is quoted or otherwise used, reference to the author’s name and the source of quotation is required.

Search all student theses