arXiv:2504.00186cs.LGcs.AI2025-04被引 4

现有鲁棒性评估基准因设计缺陷,无法真实检验模型对虚假相关性的依赖。

Are Domain Generalization Benchmarks with Accuracy on the Line Misspecified?

  • 提出判断分布偏移能否暴露模型对虚假特征依赖的充要条件
  • 发现多数主流数据集仍存在'准确率在线'现象,说明评估不严谨
  • 总结出符合自然干预场景的设计原则,指导未来基准构建

虚假相关性——模型可利用的不稳定统计捷径——预期会降低模型在分布外(OOD)的表现。然而,在多个流行的OOD泛化基准中,基础经验风险最小化(ERM)模型往往取得最高OOD准确率。此外,域内准确率提升通常伴随域外准确率上升,即所谓“准确率在线”现象,这与虚假相关性应带来损害的预期相悖。我们证明这些现象是由于当前OOD数据集未包含真正危害域外泛化的虚假相关性变化所致,即数据集设计有误。因此,现有实践并未真正测试我们想要消除的虚假信号鲁棒性。本文贡献包括:(i) 推导出分布偏移能揭示模型对虚假特征依赖的充要条件;当条件满足时,“准确率在线”现象消失;(ii) 审查主流OOD数据集,发现多数仍呈现“准确率在线”,表明其不适合作为虚假相关性鲁棒性评估基准;(iii) 梳理少数设计合理的数据集,并总结通用设计原则,如采用自然干预场景(如疫情)的数据集,以指导未来基准建设。

原文摘要 · Abstract (English)

Spurious correlations, unstable statistical shortcuts a model can exploit, are expected to degrade performance out-of-distribution (OOD). However, across many popular OOD generalization benchmarks, vanilla empirical risk minimization (ERM) often achieves the highest OOD accuracy. Moreover, gains in in-distribution accuracy generally improve OOD accuracy, a phenomenon termed accuracy on the line, which contradicts the expected harm of spurious correlations. We show that these observations are an artifact of misspecified OOD datasets that do not include shifts in spurious correlations that harm OOD generalization, the setting they are meant to evaluate. Consequently, current practice evaluates "robustness" without truly stressing the spurious signals we seek to eliminate; our work pinpoints when that happens and how to fix it. Contributions. (i) We derive necessary and sufficient conditions for a distribution shift to reveal a model's reliance on spurious features; when these conditions hold, "accuracy on the line" disappears. (ii) We audit leading OOD datasets and find that most still display accuracy on the line, suggesting they are misspecified for evaluating robustness to spurious correlations. (iii) We catalog the few well-specified datasets and summarize generalizable design principles, such as identifying datasets of natural interventions (e.g., a pandemic), to guide future well-specified benchmarks.

域泛化基准评估虚假相关性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。