LLM反思其实只是有条件重生成,无法真正纠错。
Reflection or Re-Generation? Why LLM Revision Fails Where Human Revision Succeeds

- 设计对照实验,比较人与LLM在相同条件下的修改行为
- 客观题中信息增益近乎为零,主观题反而有害
- 适合研究AI错误修正机制或对比人类思维的学者
反思是人类改进答案的核心能力。尽管大语言模型(LLMs)被频繁提示进行‘反思’,但其效果是否接近人类尚不明确。本文提出人类-LLM反思框架(HRF),在自评、互评和跨代理设置下,以受控两阶段协议对比人类与LLM的修订表现。基于每轮迭代的交叉熵降低量的信息论分析发现,LLM反思存在两种失败模式:在有限答案空间的客观任务中,信息增益ΔI≈0,表现为无差别的重生成;在主观任务中,ΔI<0,使预测偏离目标。人类修订则在两类任务中均实现正向增益。跨代理实验表明,问题出在修订环节而非输入质量:即使高质量的人类回答也会被LLM恶化。诊断分析显示,不同任务与模型下主导子步骤各异——多选题中自检存在但主观题中薄弱,且受人工错误信号引导时部分模型表现优于基线,部分却更差。根本原因在于结构缺陷:缺乏外部信息时,自我条件化修订无法缩小对目标的不确定性,因此LLM反思应理解为条件重生成,而非真正的错误驱动修正。
原文摘要 · Abstract (English)
Reflection, the ability to revisit and revise prior reasoning, is central to how humans improve their answers. Large language models (LLMs) are increasingly prompted to "reflect," yet whether this resembles human revision remains unclear. We introduce the Human-LLM Reflection Framework (HRF), a controlled two-pass protocol comparing human and LLM revision under identical conditions across self-, peer-, and cross-agent settings. Using an information-theoretic analysis based on per-iteration cross-entropy reduction, we find two failure modes of LLM reflection. On objective tasks with finite answer spaces, reflection yields near-zero information gain (Delta I approx 0), behaving as neutral re-generation indistinguishable from re-sampling. On subjective tasks, it yields significant negative gain (Delta I < 0), moving predictions away from the target. Human revision, by contrast, yields positive gain in both settings. Cross-agent experiments localize the failure to the revision step, not input quality: LLMs degrade even high-quality human responses. Diagnostic analyses (revision conditioned on first-pass correctness, and oracle-guided revision against a random-reshuffle baseline) show that which sub-step dominates varies by task and by model rather than reducing to a single mechanism: self-error detection is present on objective multiple-choice tasks but weak on subjective ones, and recovery under an oracle error signal exceeds the baseline for some models and falls below it for others. The unifying account is structural: without external information, self-conditioned revision cannot reduce uncertainty about the target, so LLM reflection is better understood as conditioned re-generation than as genuine error-driven revision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。