测试大模型推理能力,发现能察觉异常却难定位原因。
LLMs Can See the Smoke but not the Fire: Evaluating Abductive Reasoning with Elenchos

- 用形式系统变异检测任务,把推理当逆问题求解。
- 多数模型能发现系统被改但无法准确找出修改规则。
- 多处改动时模型常遗漏部分错误,适合评估推理可靠性。
大型语言模型(LLMs)在模式识别和文本生成方面表现优异,但其在反事实推理——即推断解释观察行为的潜在假设——方面的能力仍不清晰。本文提出Elenchos(源自苏格拉底式诘问法),一个生成式评估框架,将反事实推理视为结构逆问题。给定参考形式系统(如λ-演算)及其可能被修改的对应版本,代理需判断是否存在变异,并推断导致行为差异的规则修改。对前沿与中等水平的LLMs进行评估发现存在一致的“检测-归因分离”现象:模型常能察觉系统被更改,却难以确定引发差异的具体修改。在多重相互作用的变异下,性能显著下降,模型通常仅恢复出部分底层变异。初步证据表明,增加推理时间资源带来的收益递减,仅带来小幅提升,但该结论仍需进一步验证。
原文摘要 · Abstract (English)
Large language models (LLMs) excel at pattern recognition and text generation, but their capacity for abductive inference - inferring latent hypotheses that explain observed behavior - remains poorly understood. Here, we introduce Elenchos (named after the Socratic method of cross-examination), a generative evaluation framework that measures abductive reasoning as a structural inverse problem. Given a reference formal system, such as the lambda-calculus, and a potentially mutated counterpart, agents must determine whether a mutation has occurred and infer the rule modifications responsible for the resulting behavioral differences. Evaluating frontier and mid-tier LLMs reveals a consistent detection-attribution dissociation: models often recognize that a system has been altered but struggle to identify the latent mutations causing the observed discrepancies. Performance degrades substantially under interacting mutations, where models frequently recover only a subset of the underlying mutations. Preliminary evidence also suggests diminishing returns from increased inference-time reasoning, with only modest improvements under larger reasoning budgets, though this finding requires further validation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。