基于训练动态自动划分NLI测试集难易度,提升评估真实性。
How Hard is this Test Set? NLI Characterization by Exploiting Training Dynamics
- 利用模型训练过程中的动态特征,自动区分测试样本难度。
- 高难度样本上模型表现显著下降,更贴近真实语言现象。
- 仅用部分数据训练即达全量数据效果,适合模型高效评估。
自然语言推理(NLI)评估对衡量语言理解模型至关重要,但主流数据集存在系统性虚假相关性,人为抬高模型表现。为此,我们提出一种无需人工构造虚假样本的自动化方法,通过分析训练动态将常见NLI数据集的测试集划分为三个难度等级。该分类显著降低虚假相关性度量,高难度样本上模型性能明显下降,涵盖更真实多样的语言现象。将此方法应用于训练集时,仅使用部分数据训练的模型即可达到全量数据训练的效果,优于其他数据集表征技术。研究解决了NLI数据集构建的局限性,为模型性能提供了更真实的评估,对多种自然语言理解应用具有重要意义。
原文摘要 · Abstract (English)
Natural Language Inference (NLI) evaluation is crucial for assessing language understanding models; however, popular datasets suffer from systematic spurious correlations that artificially inflate actual model performance. To address this, we propose a method for the automated creation of a challenging test set without relying on the manual construction of artificial and unrealistic examples. We categorize the test set of popular NLI datasets into three difficulty levels by leveraging methods that exploit training dynamics. This categorization significantly reduces spurious correlation measures, with examples labeled as having the highest difficulty showing markedly decreased performance and encompassing more realistic and diverse linguistic phenomena. When our characterization method is applied to the training set, models trained with only a fraction of the data achieve comparable performance to those trained on the full dataset, surpassing other dataset characterization techniques. Our research addresses limitations in NLI dataset construction, providing a more authentic evaluation of model performance with implications for diverse NLU applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。