arXiv:2505.22318cs.CLcs.LG2025-05被引 1

让大模型在违背常识的假设世界中推理,发现它们常因认知冲突出错。

Flying Pigs, FaR and Beyond: Evaluating LLM Reasoning in Counterfactual Worlds

  • 引入新基准CounterLogic,分离逻辑与知识判断
  • 11个模型在反事实场景下准确率平均下降14%
  • 先标记知识冲突再推理,性能差距缩小至7%

大语言模型在面对与既有知识相悖的假设情境时,能否进行有效推理,是其推理能力的核心挑战。本文通过设计专门用于区分逻辑有效性与知识一致性的基准CounterLogic,系统评估了11个主流大模型在六种不同推理数据集上的表现。结果表明,在反事实场景下,模型平均准确率比知识一致场景下降14%。我们提出,这一差距并非源于逻辑处理缺陷,而是由上下文与参数化知识间的认知冲突所致。受人类元认知启发,我们提出一种简单有效的干预方法Flag & Reason(FaR):先提示模型识别潜在的知识冲突,再进行推理。该方法显著改善模型表现,将性能差距缩小至7%,整体准确率提升4%。研究揭示了现代大模型在复杂推理中的关键短板,并证明元认知意识可增强其鲁棒性与可靠性。

原文摘要 · Abstract (English)

A fundamental challenge in reasoning is navigating hypothetical, counterfactual worlds where logic may conflict with ingrained knowledge. We investigate this frontier for Large Language Models (LLMs) by asking: Can LLMs reason logically when the context contradicts their parametric knowledge? To facilitate a systematic analysis, we first introduce CounterLogic, a benchmark specifically designed to disentangle logical validity from knowledge alignment. Evaluation of 11 LLMs across six diverse reasoning datasets reveals a consistent failure: model accuracy plummets by an average of 14% in counterfactual scenarios compared to knowledge-aligned ones. We hypothesize that this gap stems not from a flaw in logical processing, but from an inability to manage the cognitive conflict between context and knowledge. Inspired by human metacognition, we propose a simple yet powerful intervention: Flag & Reason (FaR), where models are first prompted to flag potential knowledge conflicts before they reason. This metacognitive step is highly effective, narrowing the performance gap to just 7% and increasing overall accuracy by 4%. Our findings diagnose and study a critical limitation in modern LLMs' reasoning and demonstrate how metacognitive awareness can make them more robust and reliable thinkers.

大模型推理反事实推理元认知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。