提出检测数据集异常影响力的新方法,可判断影响是否超出随机波动范围。
Testing Most Influential Sets
- 基于线性最小二乘法推导出精确影响公式,识别极端影响分布规律。
- 发现常量大小集合在重尾数据下呈弗雷歇分布,增长集合则服从古贝尔分布。
- 可在经济、生物和机器学习中验证结论可靠性,替代主观判断。
少量关键数据点可能显著改变模型结论,甚至推翻核心发现。尽管已有研究识别出最具影响力的子集,但尚无正式方法判断最大影响力是否超过自然随机采样下的预期。本文针对线性最小二乘问题,推导出精确的影响公式,并确定最大影响的极值分布:固定大小集合在重尾数据下服从弗雷歇分布,而增长集合或轻尾情况下服从古贝尔分布。这一理论支持严格的假设检验,以判断是否存在过度影响。我们在经济学、生物学及机器学习基准测试中验证该方法,解决了争议性结论,用严谨推断取代了经验性启发式手段。
原文摘要 · Abstract (English)
Small influential data subsets can dramatically impact model conclusions, with a few data points overturning key findings. While recent work identifies these most influential sets, there is no formal way to tell when maximum influence is excessive rather than expected under natural random sampling variation. We address this gap by developing a principled framework for most influential sets. Focusing on linear least-squares, we derive a convenient exact influence formula and identify the extreme value distributions of maximal influence - the heavy-tailed Fréchet for constant-size sets and heavy-tailed data, and the well-behaved Gumbel for growing sets or light tails. This allows us to conduct rigorous hypothesis tests for excessive influence. We demonstrate through applications across economics, biology, and machine learning benchmarks, resolving contested findings and replacing ad-hoc heuristics with rigorous inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。