提出高效算法,验证线性回归对删样本的鲁棒性。
Robustness Auditing for Linear Regression: To Singularity and Beyond
- 设计可扩展算法,检验删除少量样本是否改变回归结论
- 在数百维、数万样本数据集上首次获得非平凡鲁棒性证明
- 适用于高维经济计量学数据,揭示经典研究脆弱性
近期发现,许多重要经济学研究的结论仅因移除极小比例样本(常低于0.5%)即被推翻。这些结论通常基于普通最小二乘法(OLS)回归,引发关键问题:给定数据集,能否证明某次OLS拟合对特定数量样本删除的鲁棒性?暴力方法在小型数据集上即失效;现有方法或仅能寻找候选删样子集而无法认证其不存在 [BGM20, KZC21],或在低维外计算不可行 [MR22],或需强分布假设且样本量过大而无法实用 [BP21, FH23]。本文提出一种高效算法,用于认证线性回归对样本删除的鲁棒性。我们在多个高维(4维及以上)、含数万样本的经济学经典数据集上实现并运行该算法,首次为维度≥4的数据集提供非平凡的鲁棒性证书。在数据分布假设下,证明所产边界至多相差一个1+o(1)倍因子,具有渐近紧性。
原文摘要 · Abstract (English)
It has recently been discovered that the conclusions of many highly influential econometrics studies can be overturned by removing a very small fraction of their samples (often less than $0.5\%$). These conclusions are typically based on the results of one or more Ordinary Least Squares (OLS) regressions, raising the question: given a dataset, can we certify the robustness of an OLS fit on this dataset to the removal of a given number of samples? Brute-force techniques quickly break down even on small datasets. Existing approaches which go beyond brute force either can only find candidate small subsets to remove (but cannot certify their non-existence) [BGM20, KZC21], are computationally intractable beyond low dimensional settings [MR22], or require very strong assumptions on the data distribution and too many samples to give reasonable bounds in practice [BP21, FH23]. We present an efficient algorithm for certifying the robustness of linear regressions to removals of samples. We implement our algorithm and run it on several landmark econometrics datasets with hundreds of dimensions and tens of thousands of samples, giving the first non-trivial certificates of robustness to sample removal for datasets of dimension $4$ or greater. We prove that under distributional assumptions on a dataset, the bounds produced by our algorithm are tight up to a $1 + o(1)$ multiplicative factor.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。