arXiv:2510.24295cs.CL2025-10

提出MERGE测试,评估模型对语义微变的泛化能力

MERGE: Minimal Expression-Replacement GEneralization Test for Natural Language Inference

  • 用掩码语言模型生成语义不变的表达替换变体
  • 主流模型在微小形式变化下准确率大幅下降
  • 适合关注模型泛化性与鲁棒性的研究者

随着诸多评测基准趋于饱和,构建能评估当前先进模型推理泛化能力的新数据集变得愈发重要。然而,高质量推理数据集的设计极具挑战:人工构造成本高,自动生成又不可靠,常导致覆盖范围有限的合成数据。本文提出最小表达替换泛化测试(MERGE),评估推理模型对现有评测数据集非对抗性变体的鲁棒性。通过最小表达替换(MERE)生成方法,利用掩码语言模型(MLMs)与防护过滤器自动生成高质量变体。将MERGE测试应用于自然语言推理(NLI)任务,在两个广泛使用的原始数据集上生成新NLI数据集,并用于评估多个强模型。结果表明,大语言模型与微调后的NLI模型泛化能力均较差:对形式和推理仅作微小改变的样本难以保持一致且正确的分类表现。此外,还分析了词类与源MLMs等生成因素对模型性能的影响。

原文摘要 · Abstract (English)

As many benchmarks have become saturated, it has become increasingly important to create new datasets that evaluate the generalization capacity of current state-of-the-art models in reasoning. However, designing high-quality reasoning datasets is challenging, as their manual construction is costly, and their automatic generation is unreliable, often leading to synthetic data with limited scope. In this paper, we propose the Minimal Expression-Replacement GEneralization (MERGE) test that evaluates the robustness of reasoning models against non-adversarial variants of existing evaluation datasets. We automatically obtain high-quality variants from the original instances with Minimal Expression REplacement (MERE) generation, which uses Masked Language Models (MLMs) and safeguarding filters. We apply the MERGE test to Natural Language Inference (NLI), a popular task of reasoning. We generate new NLI datasets from two widely used existing ones with the MERE generation and use them to evaluate multiple strong NLI models. The results indicate that both LLMs and fine-tuned NLI models generalize poorly: they struggle to consistently and correctly classify variants minimally different in form and reasoning from the original ones. Further, we also analyze how certain aspects in variant generation, such as the word class and the source MLMs, affect model performance.

自然语言推理泛化能力模型鲁棒性数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。