arXiv:2503.22395cs.CL2025-03被引 8

提出新数据集,揭示大模型在否定句理解上的语言差异与改进路径。

Negation: A Pink Elephant in the Large Language Models' Room?

  • 构建多语言否定蕴含数据集,分析模型处理否定的根因。
  • 发现模型规模增大可提升否定处理能力,且英语优于德语/捷克语。
  • 强调前提长度与显式程度影响模型鲁棒性,适合多语言推理研究者。

否定是决定句子语义的关键,对逻辑推理至关重要。然而,否定对大型语言模型(LLMs)仍构成重大挑战,且研究不足。我们构建并发布了两个新的文本蕴含数据集——NoFEVER-ML 和 NoSNLI-ML,涵盖英语、捷克语、德语和乌克兰语,包含含否定的样本,用于探究否定问题的根源及其表现。与以往研究不同,我们发现增加模型规模可提升其处理否定的能力。此外,模型的推理准确率与抗否定鲁棒性具有语言依赖性,前提长度与显式程度也显著影响鲁棒性。在顺序固定的投射语言(如英语)中表现优于非投射语言(如德语或捷克语)。这些蕴含数据集为解释和解决否定问题、减少大模型幻觉、提升多语言场景下的推理能力提供了基础。数据集已公开,推动后续研究。

原文摘要 · Abstract (English)

Negations are key to determining sentence meaning, making them essential for logical reasoning. Despite their importance, negations pose a substantial challenge for large language models (LLMs) and remain underexplored. We constructed and published two new textual entailment datasets NoFEVER-ML and NoSNLI-ML in four languages (English, Czech, German, and Ukrainian) with examples differing in negation. It allows investigation of the root causes of the negation problem and its exemplification: how popular LLM model properties and language impact their inability to handle negation correctly. Contrary to previous work, we show that increasing the model size may improve the models' ability to handle negations. Furthermore, we find that both the models' reasoning accuracy and robustness to negation are language-dependent and that the length and explicitness of the premise have an impact on robustness. There is better accuracy in projective language with fixed order, such as English, than in non-projective ones, such as German or Czech. Our entailment datasets pave the way to further research for explanation and exemplification of the negation problem, minimization of LLM hallucinations, and improvement of LLM reasoning in multilingual settings.

否定理解多语言推理增强数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。