arXiv:2506.11930cs.CL2025-06NeurIPS被引 8

大模型接收外部反馈后仍难完全改正错误,存在反馈阻力现象。

Feedback Friction: LLMs Struggle to Fully Incorporate External Feedback

  • 设计闭环实验环境,让模型先答题再获近完美反馈并重答。
  • 即使在理想条件下,模型仍普遍抗拒反馈,改进有限。
  • 高置信度预测更难被纠正,可用语义熵衡量反馈抵抗性。

近期研究显示大语言模型具备一定能力,在获得外部反馈后可改进回答。然而,这些模型能否有效且彻底地吸收外部反馈仍不明确。在理想情况下,若模型接收到近乎完整且准确的反馈,应能完全整合信息并得出正确答案。本文通过构建受控实验环境系统研究该问题:每个任务中,求解模型先尝试解答,随后由可访问接近完整真实答案的反馈生成器提供针对性反馈,求解模型再次尝试。我们在数学推理、知识推理、科学推理及多领域通用评估任务上测试了包括Claude 3.7(带扩展思维)在内的多种前沿语言模型。令人意外的是,即便在近理想条件下,求解模型仍持续表现出对反馈的抵抗性,我们将其称为‘反馈摩擦’。为缓解此问题,我们尝试了渐进式温度提升和显式拒绝先前错误答案等采样策略,虽有一定改善,但仍无法使模型达到目标性能。我们分析发现,模型在特定问题上的置信度(以语义熵衡量)可预测其对反馈的抵抗程度:高置信度预测更难以被外部修正。我们希望揭示这一问题,推动未来自改进研究的发展。

原文摘要 · Abstract (English)

Recent studies have shown LLMs possess some ability to improve their responses when given external feedback. However, it remains unclear how effectively and thoroughly these models can incorporate extrinsic feedback. In an ideal scenario, if LLMs receive near-perfect and complete feedback, we would expect them to fully integrate the feedback and reach correct solutions. In this paper, we systematically investigate LLMs' ability to incorporate feedback by designing a controlled experimental environment. For each problem, a solver model attempts a solution, then a feedback generator with access to near-complete ground-truth answers produces targeted feedback, after which the solver tries again. We evaluate this pipeline across a diverse range of tasks, including math reasoning, knowledge reasoning, scientific reasoning, and general multi-domain evaluations with state-of-the-art language models including Claude 3.7 with extended thinking. Surprisingly, even under these near-ideal conditions, solver models consistently show resistance to feedback, a limitation that we term Feedback Friction. To mitigate this limitation, we experiment with sampling-based strategies like progressive temperature increases and explicit rejection of previously attempted incorrect answers, which yield improvements but still fail to help models achieve target performance. We analyze Feedback Friction and find that models' confidence on specific questions, measured by semantic entropy, predicts feedback resistance: high-confidence predictions remain resistant to external correction. We hope that highlighting this issue in LLMs will help future research in self-improvement.

大模型反馈机制自改进认知偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。