2026/2027





Непараметрическая теория и методы анализа данных
Статус:
Маго-лего
Где читается:
Факультет социальных наук
Охват аудитории:
для своего кампуса
Преподаватели:
Пашков Станислав Георгиевич
Язык:
русский
Кредиты:
6
Контактные часы:
40
Программа дисциплины
Аннотация
This course is devoted to a separate section of statistical theory dealing with non-parametric methods of statistical data analysis (non-parametric statistics, NPS). This section of statistics is very often used in conjunction with more “classical” approaches based on Gaussian statistics, but it is arranged differently and requires a special approach to understanding and interpretation. Currently, more and more business decisions are made on the basis of data measured in categorical and rank scales, and therefore the relevance of this type of data analysis is increasing. Throughout the course, students will receive a theoretical and practical understanding of how to approach the procedure of non-parametric data analysis, on what types of data it is possible to do this, what needs to be considered and how to interpret the data. A special place in the course is occupied by rank regression and loglinear regression as special cases of working with data that do not have a Gaussian distribution. All the learning process is based on R language with special libraries.
Цель освоения дисциплины
- Get an idea of what non-parametric statistics is and how it differs in the process of data analysis;
- Formulate a typical algorithm of actions for diagnosing data for the need for non-parametric statistics;
- Get a comprehensive understanding of the models for setting tasks for comparing and evaluating the effects of various factors on processes that are of a categorical nature;
- Master principles of data interpretation and analytical conclusions with applications to business-related data.
Планируемые результаты обучения
- Able to fit a logistic regression model on a given dataset
- Able to use R programming language for complex statistical computations
- Can test parametric and nonparametric hypotheses
- - Implement causal inference methods (matching, instrumental variables, regression discontinuity, difference-in-difference, fixed effects) - Identify which causal assumptions are necessary for each type of statistical method
- Become familiar with non-parametric statistics
- Students become familiar with the data loading process and EDA principles.
- Students are able to use different R packages for visualization of the distribution of data, as well as the interpretation of the received data.
- Students are able to use statistical functions for testing variables for the presence of an abnormal distribution and associated key characteristics, in order to subsequently choose a fundamentally different stack of functions and methods for further statistical data analysis
- Students get acquainted with special packages pdfCluster, BayesBinMix, functions suitable with DBSCAN approaches (clues, base-R).
- Students are introduced to the theoretical background and challenges of working with non-parametric data in order to obtain statistically valid inferences and interpretations
Содержание учебной дисциплины
- Section 1: Course Structure. Types of data for NPS/EDA. The EDA framework
- Section 2: Statistical Distribution (Part 1): Theoretical Genesis.
- Section 3: Statistical Distribution (Part 2): Applied Principles.
- Section 4: NPS Cluster Analsysis: General Framework and Genesis
- Section 5: Non-Parametric Regression Adventures (Part 1): The Genesis of NP-reg and ordinal/nominal models
- Sections 6: Non-Parametric Regression Adventures (Part 2). Log-Linear Models. The Genesis of Contingency Tables
- Section 7: Principles of non-parametric Times-Series analysis
Элементы контроля
- Practice Task 1This laboratory assignment consolidates theoretical knowledge of non-parametric methods through hands-on analysis of two real-world datasets. Students practice: (1) scale-aware data preparation and exploratory diagnostics using modern EDA packages; (2) kernel density estimation with principled bandwidth selection; (3) exact versus asymptotic inference in rank-based tests; (4) score-function frameworks for robust group comparison; and (5) reproducible research practices. Both code quality and statistical interpretation are assessed. The assignment consists of 5 mandatory problems (10 points total) plus optional challenge tasks (up to +1 bonus point).
- Practice Task 2This laboratory assignment consolidates theoretical knowledge of non-parametric clustering methods through hands-on analysis of the dataset given. Students practice: (1) scale-aware data preparation and exploratory diagnostics; (2) distribution visualization, non-parametric testing, and correlation-based variable selection; (3) application of density-based, median-based, and correlation-based clustering techniques (pdfCluster, Gmedian, DBSCAN, t-SNE); and (4) reproducible research practices. The assignment uses a single dataset for all students to ensure consistency and efficient grading. It consists of 4 mandatory problems (10 points total) plus optional challenge tasks (up to +1 bonus point).
- Practice Task 3This laboratory assignment focuses on the main types of non-parametric regression analysis using the dataset given. Students build progressively more sophisticated models: from hierarchical linear regression through robust rank-based methods (Kendall–Theil Sen, Siegel, Rfit) and quantile regression, to Generalized Additive Models (GAMs). The assignment uses a single dataset to ensure comparability across submissions and consists of 6 problems (10 points total). Both code production quality and algorithm interpretation are assessed.
- Final ProjectThe Final Project is the capstone assessment of the course, requiring each student to design and execute a complete full-cycle research study that integrates all non-parametric methods covered throughout the module: exploratory data analysis with kernel density exploration, normality testing, non-parametric clustering, non-parametric regression (Kendall–Theil, Siegel, rank-based, quantile, GAMs), and optionally log-linear analysis or tree-based ensemble methods (RandomForest/CatBoost). Unlike the laboratory assignments (which use instructor-provided datasets), the Final Project demands independent research design: students must formulate a theoretically grounded research question, select and justify a data source (Kaggle, GEM, ESS, WVS, or equivalent), construct a theoretical argumentation framework linked to peer-reviewed literature, execute a rigorous non-parametric analytical pipeline in R (with Python permitted upon explicit justification of package analogues), and articulate findings in both a comprehensive Jupyter/RMarkdown report and a concise defense presentation (maximum 10 content slides + 2 technical slides). The project may be completed individually or in pairs (with doubled requirements for pairs). The assessment splits into two weighted components: Report Quality (70%) and Small Presentation + Defense (30%). This form of control develops the complete competency profile of an applied non-parametric researcher — from problem formulation through empirical execution to scientific communication — and produces a portfolio-grade artifact suitable for thesis development or industry case-study portfolios.
Промежуточная аттестация
- 2026/2027 4th module0.15 * Practice Task 1 + 0.25 * Practice Task 3 + 0.35 * Final Project + 0.25 * Practice Task 2
Список литературы
Рекомендуемая основная литература
- 9781292034898 - Agresti, Alan; Finlay, Barbara - Statistical Methods for the Social Sciences - 2014 - Pearson - https://search.ebscohost.com/login.aspx?direct=true&db=nlebk&AN=1418314 - nlebk - 1418314
- Agresti, A. (2013). Categorical Data Analysis (Vol. Third edition). Hoboken, NJ: Wiley. Retrieved from http://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsebk&AN=769330
- Agresti, A. (2015). Foundations of Linear and Generalized Linear Models. Hoboken, New Jersey: Wiley. Retrieved from http://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsebk&AN=941245
- Agresti, A. (2017). Statistics: The Art and Science of Learning From Data, Global Edition. Pearson.
- Bruce E. Hansen, Donald W. K. Andrews, A. Ronald, Gallant Douglas, W. Nychka, & James G. Mackinnon. (n.d.). Semi-Nonparametric Maximum Likelihood Estimation. Retrieved from http://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsbas&AN=edsbas.7BD2F74E
- Corder, G. W., & Foreman, D. I. (2014). Nonparametric Statistics : A Step-by-Step Approach (Vol. Second edition). Hoboken, New Jersey: Wiley. Retrieved from http://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsebk&AN=798830
- Francois Treves. (2013). Topological Vector Spaces, Distributions and Kernels. [N.p.]: Dover Publications. Retrieved from http://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsebk&AN=1151250
- Härdle, W., Müller, M., Sperlich, S. A., & Werwatz, A. (2004). Nonparametric and Semiparametric Models. Switzerland, Europe: Springer. Retrieved from http://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsbas&AN=edsbas.121C8F13
- Myatt, G. J., & Johnson, W. P. (2014). Making Sense of Data I : A Practical Guide to Exploratory Data Analysis and Data Mining (Vol. Second edition). Hoboken, New Jersey: Wiley. Retrieved from http://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsebk&AN=809795
- Wasserman, L. All of nonparametric statistics. – Springer Science & Business Media, 2006. – 270 pp.
Рекомендуемая дополнительная литература
- Wickham, H., & Grolemund, G. (2016). R for Data Science : Import, Tidy, Transform, Visualize, and Model Data (Vol. First edition). Sebastopol, CA: Reilly - O’Reilly Media. Retrieved from http://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsebk&AN=1440131