发现推理链末尾多余推理会损害模型训练效果
Diagnosing Harmful Continuation in Answer-Correct Long-CoT Training Traces

- 通过删除答案后冗余推理,提升微调效果
- 发现末尾推理存在局部不确定性与方向退化
- 提出轻量级边界检测方法,适合优化长推理数据
长链式思维(CoT)广泛用于推理类大模型的监督微调,但即使答案正确,不同推理轨迹仍会导致显著不同的微调结果。本文研究答案正确轨迹中的结论后延续现象:即答案已充分支持,但推理继续进行且仍被纳入监督目标。为检验其影响,使用仅删除编辑器构造答案保持的后缀移除,并对比原始与处理后轨迹的CoT微调效果。结果显示,移除该延续后微调性能提升,表明此类延续有害。进一步通过不确定性和隐藏状态进展分析,发现持续的局部不确定性与减弱的终端方向性进展形成不匹配。最后提出轻量级边界代理方法HCC,近似识别该延续边界。
原文摘要 · Abstract (English)
Long chain-of-thought (CoT) traces are widely used as supervision for reasoning-oriented LLM SFT, yet answer-correct traces can still lead to markedly different fine-tuning outcomes. We study post-conclusion continuation in answer-correct long-CoT data: a continuation where the answer appears sufficiently supported, but the trace continues with additional reasoning that remains in the supervised target. To test its training effect, we use a delete-only editor to construct answer-preserving suffix removal and compare CoT-based SFT on the original and processed traces. We observe improved SFT outcomes after removing the editor-identified post-conclusion continuation, suggesting that this continuation is harmful to training in our setting. We therefore refer to this empirically supported phenomenon as harmful continuation. Beyond this intervention, we further characterize the removed post-conclusion continuation through uncertainty and hidden-state progress. We observe persistent local uncertainty together with weakened terminal-directional progress, forming an uncertainty--geometry mismatch. Finally, we instantiate Harmful Continuation Cut (HCC), a lightweight boundary proxy that approximates the editor-identified post-conclusion continuation boundary.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。