研究推理模型如何发现并纠正错误思路,发现其自省能力有限。
How Well Can Reasoning Models Identify and Recover from Unhelpful Thoughts?
- 通过四种错误思维类型测试模型自省能力
- 大模型反而更难从干扰思路中恢复,性能显著下降
- 小模型抗干扰更强,适合构建更安全的推理系统
近期推理模型展现出反思、回溯和自我验证的能力,这对发现错误、得出正确答案至关重要。本文探讨模型在识别和纠正四类无帮助思维方面的有效性:无信息冗余、与问题无关、偏离原问题、导致错误答案的思维。结果表明,模型能有效识别多数错误思维,但一旦这些思维被注入推理过程,便难以恢复,导致性能明显下降。模型倾向于盲目延续干扰性思路,说明其自省能力并非普遍意义上的元认知。此外,观察到非/反尺度效应:大模型比小模型更难从简短无关思维中恢复,即使被指令重新评估。通过注入无关思维的越狱实验表明,最小模型受有害提示影响最小。研究呼吁改进推理模型的自我评估机制,以提升推理能力和系统安全性。
原文摘要 · Abstract (English)
Recent reasoning models show the ability to reflect, backtrack, and self-validate their reasoning, which is crucial in spotting mistakes and arriving at accurate solutions. A natural question that arises is how effectively models can perform such self-reevaluation. We tackle this question by investigating how well reasoning models identify and recover from four types of unhelpful thoughts: uninformative rambling thoughts, thoughts irrelevant to the question, thoughts misdirecting the question as a slightly different question, and thoughts that lead to incorrect answers. We show that models are effective at identifying most unhelpful thoughts but struggle to recover from the same thoughts when these are injected into their thinking process, causing significant performance drops. Models tend to naively continue the line of reasoning of the injected irrelevant thoughts, which showcases that their self-reevaluation abilities are far from a general "meta-cognitive" awareness. Moreover, we observe non/inverse-scaling trends, where larger models struggle more than smaller ones to recover from short irrelevant thoughts, even when instructed to reevaluate their reasoning. We demonstrate the implications of these findings with a jailbreak experiment using irrelevant thought injection, showing that the smallest models are the least distracted by harmful-response-triggering thoughts. Overall, our findings call for improvement in self-reevaluation of reasoning models to develop better reasoning and safer systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。