首次系统梳理大模型的反事实推理,厘清生成与选择两个关键阶段。
Wiring the 'Why': A Unified Taxonomy and Survey of Abductive Reasoning in LLMs

- 提出生成-选择双阶段统一定义,拆解反事实推理过程。
- 构建涵盖任务、数据集、方法的完整分类体系。
- 适合研究大模型推理能力或认知机制的学者参考。
尽管反事实推理在人类发现与理解中具有基础性作用,但在大语言模型(LLMs)中仍相对未被充分探索。本文首次系统综述了大模型中的反事实推理,从哲学基础追溯到当代人工智能实现。为解决领域内概念混淆和任务定义分散的问题,我们提出一个统一的两阶段定义:假设生成(模型填补认知空白,提出候选解释)与假设选择(评估候选并选出最合理解释)。基于此,我们构建了全面的文献分类体系,按任务类型、数据集、方法论和评估策略对已有工作进行归类。为实证验证框架,我们对当前主流大模型在反事实任务上开展小型基准测试,并对比不同模型规模、模型家族、评估方式及生成/选择任务范式的性能差异。结合最新实证结果,分析反事实推理与演绎、归纳推理的关系,揭示大模型整体推理能力的内在关联。分析发现当前方法存在诸多缺陷,包括基准设计僵化、领域覆盖狭窄、训练框架单一以及对反事实推理机制理解不足等。
原文摘要 · Abstract (English)
Regardless of its foundational role in human discovery and sense-making, abductive reasoning--the inference of the most plausible explanation for an observation--has been relatively underexplored in Large Language Models (LLMs). Despite the rapid advancement of LLMs, the exploration of abductive reasoning and its diverse facets has thus far been disjointed rather than cohesive. This paper presents the first survey of abductive reasoning in LLMs, tracing its trajectory from philosophical foundations to contemporary AI implementations. To address the widespread conceptual confusion and disjointed task definitions prevalent in the field, we establish a unified two-stage definition that formally categorizes prior work. This definition disentangles abduction into Hypothesis Generation, where models bridge epistemic gaps to produce candidate explanations, and Hypothesis Selection, where the generated candidates are evaluated and the most plausible explanation is chosen. Building upon this foundation, we present a comprehensive taxonomy of the literature, categorizing prior work based on their abductive tasks, datasets, underlying methodologies, and evaluation strategies. In order to ground our framework empirically, we conduct a compact benchmark study of current LLMs on abductive tasks, together with targeted comparative analyses across model sizes, model families, evaluation styles, and the distinct generation-versus-selection task typologies. Moreover, by synthesizing recent empirical results, we examine how LLM performance on abductive reasoning relates to deductive and inductive tasks, providing insights into their broader reasoning capabilities. Our analysis reveals critical gaps in current approaches--from static benchmark design and narrow domain coverage to narrow training frameworks and limited mechanistic understanding of abductive processes...
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。