审视虚假相关性评测基准的可靠性,提出选择方法的实用指南。
Reassessing the Validity of Spurious Correlations Benchmarks
- 定义评测基准应满足的三项核心标准
- 发现现有基准间结果差异大,部分无效
- 给出基于相似基准选方法的实操建议
神经网络在数据存在虚假相关性时可能失效。为理解此现象,研究者提出了多个虚假相关性评测基准以评估缓解方法。然而我们发现这些基准存在显著分歧:在某一基准表现优异的方法在另一基准上表现不佳。本文探讨了这种分歧,通过定义三项理想基准应具备的标准来评估其有效性。结果表明,某些基准并非衡量方法性能的可靠指标,且多种缓解方法也未达到广泛使用的鲁棒性要求。最后,本文为从业者提供了一个简单实用的方案:根据与自身问题最相似的基准来选择方法。
原文摘要 · Abstract (English)
Neural networks can fail when the data contains spurious correlations. To understand this phenomenon, researchers have proposed numerous spurious correlations benchmarks upon which to evaluate mitigation methods. However, we observe that these benchmarks exhibit substantial disagreement, with the best methods on one benchmark performing poorly on another. We explore this disagreement, and examine benchmark validity by defining three desiderata that a benchmark should satisfy in order to meaningfully evaluate methods. Our results have implications for both benchmarks and mitigations: we find that certain benchmarks are not meaningful measures of method performance, and that several methods are not sufficiently robust for widespread use. We present a simple recipe for practitioners to choose methods using the most similar benchmark to their given problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。