arXiv:2412.15269cs.CLcs.AI2024-12中稿 · publication at the…被引 1

发现语言模型校准越好,越可能依赖捷径,反而不可靠。

The Reliability Paradox: Exploring How Shortcut Learning Undermines Language Model Calibration

  • 通过分析校准误差与捷径学习关系,揭示模型可靠性陷阱。
  • 校准误差低的模型决策规则更不具泛化能力。
  • 提醒研究者警惕表面可靠的模型,适合关注模型可信度的读者。

预训练语言模型(PLMs)在自然语言处理中取得显著进展,但近期研究发现其存在校准问题,即模型置信度估计不准确。现有评估方法通常认为校准误差越低,预测越可靠。然而,微调后的PLMs常采用捷径学习策略,导致过度自信的预测,看似性能提升,实则缺乏泛化能力。本文探究校准误差与捷径学习之间的关系,发现校准表现看似优异的模型,其决策规则更具非泛化性。这一结果挑战了‘校准良好即可靠’的普遍认知。研究呼吁弥合校准与泛化目标之间的差距,推动构建真正鲁棒且可靠的语言模型综合框架。

原文摘要 · Abstract (English)

The advent of pre-trained language models (PLMs) has enabled significant performance gains in the field of natural language processing. However, recent studies have found PLMs to suffer from miscalibration, indicating a lack of accuracy in the confidence estimates provided by these models. Current evaluation methods for PLM calibration often assume that lower calibration error estimates indicate more reliable predictions. However, fine-tuned PLMs often resort to shortcuts, leading to overconfident predictions that create the illusion of enhanced performance but lack generalizability in their decision rules. The relationship between PLM reliability, as measured by calibration error, and shortcut learning, has not been thoroughly explored thus far. This paper aims to investigate this relationship, studying whether lower calibration error implies reliable decision rules for a language model. Our findings reveal that models with seemingly superior calibration portray higher levels of non-generalizable decision rules. This challenges the prevailing notion that well-calibrated models are inherently reliable. Our study highlights the need to bridge the current gap between language model calibration and generalization objectives, urging the development of comprehensive frameworks to achieve truly robust and reliable language models.

语言模型校准捷径学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。