arXiv:2409.05197cs.CLcs.AI2024-09EMNLP被引 15

大模型在多跳推理中易被看似合理的干扰项误导,影响判断准确率。

Seemingly Plausible Distractors in Multi-Hop Reasoning: Are Large Language Models Attentive Readers?

  • 构建看似合理但错误的多跳推理链,测试模型真实推理能力。
  • 部分模型在干扰项下F1分数下降最高达45%。
  • 适合关注大模型可信推理与评测方法的研究者阅读。

当前先进的大型语言模型(LLMs)具备从阅读理解到复杂推理等多种能力。本文聚焦其多跳推理能力:从多个文本源中识别并整合信息。鉴于现有评测基准中存在简化线索,使模型可绕过真实推理过程,本文探究了LLMs是否倾向于利用此类线索。研究发现,模型确实会规避真正的多跳推理,且方式比之前的预训练语言模型更隐蔽。为此,我们提出一个挑战性多跳推理基准,通过生成看似合理但最终导致错误答案的推理链进行测试。评估多个开源及专有顶级LLM后发现,面对此类干扰时,其性能显著下降,F1分数最高相对降低45%。深入分析表明,尽管模型能忽略误导性词汇线索,但误导性推理路径仍构成重大挑战。

原文摘要 · Abstract (English)

State-of-the-art Large Language Models (LLMs) are accredited with an increasing number of different capabilities, ranging from reading comprehension, over advanced mathematical and reasoning skills to possessing scientific knowledge. In this paper we focus on their multi-hop reasoning capability: the ability to identify and integrate information from multiple textual sources. Given the concerns with the presence of simplifying cues in existing multi-hop reasoning benchmarks, which allow models to circumvent the reasoning requirement, we set out to investigate, whether LLMs are prone to exploiting such simplifying cues. We find evidence that they indeed circumvent the requirement to perform multi-hop reasoning, but they do so in more subtle ways than what was reported about their fine-tuned pre-trained language model (PLM) predecessors. Motivated by this finding, we propose a challenging multi-hop reasoning benchmark, by generating seemingly plausible multi-hop reasoning chains, which ultimately lead to incorrect answers. We evaluate multiple open and proprietary state-of-the-art LLMs, and find that their performance to perform multi-hop reasoning is affected, as indicated by up to 45% relative decrease in F1 score when presented with such seemingly plausible alternatives. We conduct a deeper analysis and find evidence that while LLMs tend to ignore misleading lexical cues, misleading reasoning paths indeed present a significant challenge.

大模型多跳推理评测基准推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。