Double Machine Learning for Partially Linear Models with Endogeneity and Multivariate Sample Selection
Bogdan Potanin (HSE University) has published an article titled "Double machine learning for a partially linear model with endogenous treatments and multivariate sample selection" in the journal Statistics and Computing. The paper proposes double machine learning (DML) estimators for a partially linear model with endogenous treatments and multivariate sample selection. Asymptotic normality of the estimators is proved under mild regularity conditions, and their finite sample properties are studied on simulated data. The proposed approach is extended to the case of endogenous switching and is illustrated by estimating the Engel curve using RLMS-HSE data.
Statistics and Computing has published an article by Bogdan Potanin (HSE University) titled "Double machine learning for a partially linear model with endogenous treatments and multivariate sample selection."
The paper addresses the estimation of a partially linear model in which two common problems arise simultaneously: endogeneity of the treatment variables and non-random sample selection governed by several selection equations at once. Such settings are typical for applied research: for instance, wages and labor supply are observable only for employed individuals and only if the respondent answers the corresponding survey questions. So far, the estimation of a partially linear model under multivariate sample selection has received almost no attention in the literature.
The author proposes estimators based on double machine learning (DML). To construct them, the efficient influence function is derived, a Neyman-orthogonal score is established on its basis, and the estimator itself relies on cross-fitting with a nested sample split. Multivariate selection is handled through the control function approach, in which the conditional probabilities of selection from each equation serve as additional regressors. Under mild regularity conditions, the proposed estimators are shown to be consistent and asymptotically normal. Modifications of the method are also provided for the cases of exogenous treatments and endogenous regime switching.
Simulation results demonstrate that ignoring sample selection, or reducing it to a single combined equation, leads to a substantial bias, whereas the proposed DML estimators remain accurate. When the data generating process deviates considerably from linearity, they outperform the conventional control function approach, and nested cross-fitting noticeably improves precision. In the empirical application, the method is used to estimate the Engel curve with the 2023 wave of the Russia Longitudinal Monitoring Survey (RLMS-HSE): the share of income spent on food is observable only if a household answers both the income and the food spending questions. The resulting DML estimates are close to those of the parametric model with multivariate sample selection, confirming the robustness of the findings.

