Robustness of ML models to sample removals: Theory for OLS, methods and implications

Boaz Nadler (Weizmann Institute of Science)

Abstract: For learned models to be trustworthy, it is essential to verify their robustness to perturbations in the training data. Recent studies have found that for various datasets, learned models may change significantly with the removal of even less than one percent of the training samples. This led several authors to propose a more stringent form of robustness, denoted robustness auditing. Going beyond confidence intervals and bootstrap methods, robustness auditing considers the most influential subset selection (MISS) problem: stability to the removal of any subset of k samples from the training set. In this talk, I’ll present a theoretical study of this form of robustness for ordinary least squares (OLS). We prove that under mild conditions, OLS is robust to removal of any k« n i.i.d. samples, for any dimension p < n. In contrast, if k is proportional to n, then OLS is provably non-robust. We further develop a fast method to approximate the MISS problem. Finally, we revisit prior analyses that found several datasets to be highly non-robust to sample removals. While this seems to contradict our theoretical results, we demonstrate that the sensitivity is due to either heavy-tailed responses or correlated samples.