测试大模型能否发现隐含线索,发现顶级模型仅42.8%准确。
MixRea: Benchmarking Explicit-Implicit Reasoning in Large Language Models

- 设计9类推理任务,用显性/隐性信息混合题评估模型注意力
- 21个模型平均表现仅42.8%,顶尖模型仍严重遗漏关键线索
- 提出新提示法补全潜在因果关系,提升模型推理可靠性
大型语言模型日益应用于高风险决策。受人类认知中‘无意盲视’理论启发,我们探究训练于人类偏好语料(含注意力偏差)的模型是否也存在类似缺陷:在明确任务指令下忽视细微但重要的上下文线索。为此,我们提出‘显性-隐性推理’任务,并构建包含2,246道多选题的基准测试集MixRea,覆盖9种推理类型,且显性与隐性信息分布各异。对21个先进LLM的评估显示,即使最佳模型Gemini 2.5 Pro的推理一致性也仅为42.8%,揭示普遍存在的无意盲视现象。为缓解此问题,我们提出潜在关系补全提示法(PRCP),通过恢复被忽略的因果关联来提升推理能力。进一步分析表明,该局限在多种多源推理任务中持续存在,凸显开发更符合认知规律模型的必要性。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly integrated into high-stakes decision-making. Inspired by the theory of \emph{inattentional blindness} in human cognition, we investigate whether LLMs, trained on human-preferred corpora that embed attentional biases, exhibit a similar limitation: \emph{failing to attend to subtle yet important contextual cues under explicit task instructions}. To evaluate this, we introduce the task of \textbf{explicit-implicit reasoning} and present \textbf{MixRea}, a benchmark of 2,246 multiple-choice questions across 9 reasoning types with varying distributions of explicit and implicit information. Evaluation of 21 advanced LLMs shows that even the best-performing reasoning model (Gemini 2.5 Pro) achieves only 42.8\% consistency, revealing widespread inattentional blindness. To mitigate this, we propose \textbf{Potential Relation Completion Prompting (PRCP)}, a prompting method that improves reasoning by recovering overlooked causal relations. Further analysis shows that this limitation persists across diverse multi-source reasoning tasks, highlighting the need for more cognitively aligned models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。