arXiv:2605.01704cs.CLcs.AI2026-05被引 2

大模型多步推理易陷入思维陷阱,导致推理质量下降却误保答案正确。

The Reasoning Trap: An Information-Theoretic Bound on Closed-System Multi-Step LLM Reasoning

论文配图:The Reasoning Trap: An Information-Theoretic Bound on Closed-System Multi-Step LLM Reasoning
图 1 · 摘自论文原文
  • 用信息论证明多智能体辩论会弱化推理依据,而非提升多样性。
  • 在科学事实数据集上,推理可信度下降43%,多数投票机制几乎丢尽可信度。
  • 提出基于证据的苏格拉底式追问,可恢复98%的推理可信度,适合改进AI可解释性。

当同一语言模型被多次提示进行辩论时,其输出呈现单一视角的不同表述,而非多元观点。多智能体辩论(MAD)及更广泛的封闭系统推理中,尽管答案准确率得以保持,但推理过程的质量持续退化。我们称此现象为‘辩论陷阱’与‘推理陷阱’,并提出一套以证据为基础的推理失败理论框架。该框架包含三部分:(i) SFS(支持忠实度评分),用于验证分解后原子命题与证据的一致性(分解器无关排序:斯皮尔曼相关系数=1.0);(ii) EGSR(证据基础苏格拉底推理),以证据驱动的提问替代对抗性论证;(iii) 定理1(DPI界):在标准MAD下,链式结构E→O⁰→O¹→…满足马尔可夫性,数据处理不等式表明期望信息量递减,即E[I(E;O^{t+1})] ≤ E[I(E;O^t)]。配套结果包括开放系统恢复(定理2)、EGSR累积效应(引理2)与投票聚合下限(命题1),将多步推理按其与证据的信息关系分类。在SciFact(300个主张)和FEVER(1000个主张)共16种条件下,DebateCV(C13)保留88%基准准确率,但SFS下降43%;多数投票的MAD(C15)使SFS降至基线的1.7%(p < 10⁻⁶,d = -0.96);EGSR则恢复至98%。一项包含韩语10×30和英语3×200样本的群体研究发现,跨评估者一致性(Fleiss kappa)≤ +0.018,且领域与语言间评分者内部一致性波动为0.8–1.4,表明用于校准忠实度指标的人类共识本身并不稳定。我们提出一个可检验猜想:任何维持定理1马尔可夫结构的封闭系统推理协议,在期望上都受相同DPI界约束。

原文摘要 · Abstract (English)

When copies of the same language model are prompted to debate, they produce diverse phrasings of one perspective rather than diverse perspectives. Multi-agent debate (MAD), and more broadly closed-system reasoning where agents iteratively transform each other's outputs, tends to preserve answer accuracy while degrading the reasoning behind those answers. We name the multi-agent case the Debate Trap and the broader phenomenon the Reasoning Trap, offering a programmatic theory of evidence-grounded reasoning failure.The framework has three parts: (i) SFS (Supported Faithfulness Score), a claim-level metric verifying decomposed atomic claims against provided evidence (decomposer-invariant rankings: Spearman rho=1.0); (ii) EGSR (Evidence-Grounded Socratic Reasoning), replacing adversarial argumentation with evidence-grounded inquiry; (iii) Theorem 1 (DPI Bound): under standard MAD, the chain E -> O^0 -> O^1 -> ... is Markov, and the Data Processing Inequality implies E[I(E;O^{t+1})] <= E[I(E;O^t)]. Three companion results -- open-system recovery (Theorem 2), EGSR accumulation (Lemma 2), and vote-aggregation floor (Proposition 1) -- partition multi-step LLM reasoning by its information-theoretic relationship to E. Across 16 conditions on SciFact (300 claims) and FEVER (1,000 claims), DebateCV (C13) preserves 88% of baseline accuracy while SFS drops 43%; majority-vote MAD (C15) reduces SFS to 1.7% of baseline (p < 10^{-6}, d = -0.96); EGSR recovers 98%. An R6 cohort study (Korean n=10x30 FEVER; English n=3x200 SciFact) finds inter-rater Fleiss kappa <= +0.018 with 0.8-1.4 Likert intra-rater shifts across language and domain -- the human agreement that faithfulness metrics have been calibrated against is not itself stable. We offer one falsifiable conjecture: any closed-system reasoning protocol preserving Theorem 1's Markov structure is, in expectation, subject to the same DPI bound.

大模型推理信息论可信度评估多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。