arXiv:2508.18407cs.CLcs.AI2025-08

QA模型的分布外评估可能无法发现对捷径的依赖。

Can Out-of-Distribution Evaluations Uncover Reliance on Shortcuts? A Case Study in Question Answering

  • 用多个分布外数据集测试问答模型对捷径的鲁棒性。
  • 部分评估数据集表现不如简单分布内测试,效果很差。
  • 提醒研究者谨慎使用现有分布外评估,尤其在推理时需警惕捷径依赖。

当前多数AI研究通过分布外(OOD)数据集上的表现来评估模型的泛化能力。尽管实用,这类评估基于一个关键假设:分布外评估能反映真实部署中的潜在失败。本文挑战这一假设,将分布外评估结果与已知问答模型对虚假特征或预测捷径的依赖进行对比。研究发现,不同分布外数据集对模型抗捷径能力的评估质量差异巨大,某些甚至不如简单的分布内评估。这部分归因于虚假捷径在分布内与分布外数据中共享,也发现部分数据集在训练与评估用途上严重脱节。本工作揭示了主流分布外评估在泛化性评测中的局限性,并为问答及其他领域更稳健的评估提供了方法与建议。

原文摘要 · Abstract (English)

A majority of recent work in AI assesses models' generalization capabilities through the lens of performance on out-of-distribution (OOD) datasets. Despite their practicality, such evaluations build upon a strong assumption: that OOD evaluations can capture and reflect upon possible failures in a real-world deployment. In this work, we challenge this assumption and confront the results obtained from OOD evaluations with a set of specific failure modes documented in existing question-answering (QA) models, referred to as a reliance on spurious features or prediction shortcuts. We find that different datasets used for OOD evaluations in QA provide an estimate of models' robustness to shortcuts that have a vastly different quality, some largely under-performing even a simple, in-distribution evaluation. We partially attribute this to the observation that spurious shortcuts are shared across ID+OOD datasets, but also find cases where a dataset's quality for training and evaluation is largely disconnected. Our work underlines limitations of commonly-used OOD-based evaluations of generalization, and provides methodology and recommendations for evaluating generalization within and beyond QA more robustly.

问答系统分布外评估模型可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。