模型自信不等于正确,数据难易评估需多维度考量
Your Model is Overconfident, and Other Lies We Tell Ourselves
- 用标注分歧、训练动态和模型置信度三指标衡量数据内在难度
- 29个模型在3个数据集上验证,三者关系非线性且不单调
- 揭示模型过自信的根源,适合评估与改进NLP模型的研究者
神经网络NLP模型评估中,例题固有的模糊性所导致的内在难度是一个关键却常被忽视的因素。本文通过在三个数据集上对29个模型进行综合分析,研究了标注分歧、训练动态和模型置信度等内在难度度量之间的相互作用与差异。结果表明,尽管这些度量间存在相关性,但其关系既非线性也非单调。通过解耦不确定性维度,我们旨在深化对数据复杂性的理解,并为评估与改进NLP模型提供更精准的视角。
原文摘要 · Abstract (English)
The difficulty intrinsic to a given example, rooted in its inherent ambiguity, is a key yet often overlooked factor in evaluating neural NLP models. We investigate the interplay and divergence among various metrics for assessing intrinsic difficulty, including annotator dissensus, training dynamics, and model confidence. Through a comprehensive analysis using 29 models on three datasets, we reveal that while correlations exist among these metrics, their relationships are neither linear nor monotonic. By disentangling these dimensions of uncertainty, we aim to refine our understanding of data complexity and its implications for evaluating and improving NLP models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。