模型的思考行为越像人类,越不一定更准确。
Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models

- 用行为提升度量化推理中各行为对正确率的影响。
- 自修正、假设测试等被训练强化的行为,实际与正确率关联弱。
- 真正关键的可信度校准和自我意识,反被忽略。
哪些推理行为与模型正确答案相关?推理训练是否放大了这些行为?这一区别至关重要:推理训练可能让推理过程看起来更审慎,但未必放大真正有助于正确的行为。我们提出行为提升度(Behavioral Lift)指标,衡量某一行为在推理轨迹中存在与否时正确率的变化。在15个模型和6个涵盖纯文本与图文推理的基准上,我们对15,282条推理轨迹进行标注,使用涵盖大语言模型(LLM)与视觉语言模型(VLM)的核心行为分类体系。研究发现存在‘放大-提升差距’:思考模型显著放大了自我修正、假设检验和不确定性承认,而与正确率最强相关的却是可信度校准、知识对齐和自我意识。可信度校准在两种模态中均为最强正向信号,却几乎未被放大;不确定性承认被放大3–7倍,却与正确率弱相关或负相关。结果表明,推理训练并未优先放大高提升度行为,提示应设计以校准和真实感为核心的推理目标,而非仅关注表面形式。
原文摘要 · Abstract (English)
Which reasoning behaviors are associated with correct answers in reasoning models, and does reasoning-oriented training amplify those behaviors? This distinction is important because reasoning-oriented training can make traces look more deliberative without amplifying the behaviors most tied to model correctness. We quantify this mismatch with Behavioral Lift, a metric that measures how much correctness changes when a behavior is present versus absent in a model's reasoning trace. Across 15 models and 6 benchmarks spanning text-only and vision-language reasoning, we annotate 15,282 traces with a taxonomy whose core behaviors are defined for both LLM and VLM traces. We find evidence for an Amplification-Lift Gap, in which thinking models strongly amplify self-correction, hypothesis testing, and uncertainty acknowledgment, while the highest-lift behaviors are confidence calibration, knowledge alignment, and self-awareness. Confidence calibration is among the strongest positive signals of correctness in both modalities, yet is barely amplified; uncertainty acknowledgment is amplified by 3--7$\times$, yet is weakly or negatively associated with correctness. We find that reasoning-oriented training does not preferentially amplify the highest-Lift behaviors, motivating process-level objectives that reward calibrated and grounded reasoning rather than surface form alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。