自修正能提升模型表现,但只在特定任务中有效。
When Does Intrinsic Self-Correction Help? A Task-Sensitive Analysis
- 根据任务类型设计不同自修正机制
- 在约束验证等场景下性能稳定提升
- 适合需要多轮推理的任务使用
内在自修正(SC)通过让大模型在无外部反馈的情况下重新审视自身初始回答来改进输出。近期研究质疑该方法的可靠性,指出模型常难以判断初始回答是否正确。本文从任务敏感性视角分析SC:不笼统评估其效果,而是考察其在验证显式约束、重审复杂推理过程、以及词游戏任务中对比不同策略时的表现。在多个基准和模型上,我们发现当任务结构支持上述修正机制时,SC可带来一致的性能提升。结果表明,自修正应被视为依赖任务的推理阶段策略,其有效性取决于修正环节在具体任务中的作用,而非普遍适用的输出优化方法。
原文摘要 · Abstract (English)
Intrinsic self-correction (SC) aims to improve large language model outputs by prompting a model to revisit its own initial answer without external feedback. Recent studies have questioned the reliability of this approach, showing that models often struggle to judge whether their initial responses are correct. In this work, we take a task-sensitive view of SC. Rather than asking whether it works in general, we examine settings where SC may operate through different mechanisms: verifying explicit constraints, revisiting a complex reasoning process, or providing a second opinion over competing strategies in word-game tasks. Across multiple benchmarks and models, we find that SC can yield consistent performance gains when the underlying task structure facilitates these modes of revision. These results suggest that SC is best understood as a task-dependent inference-time strategy whose usefulness depends on the role the revision stage can play in a given task, rather than as a uniformly reliable method for improving initial model outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。