发现推理模型中的话语标记词能提升准确率,但小样本训练效果仍不及大规模训练。
Oops, Wait: Discourse Tokens Matter in Reasoning Model
- 分析推理轨迹中话语标记词与答案正确性的关联性。
- 小样本微调可部分复现话语模式,使准确率提升约15%。
- 适合关注高效训练与推理机制的模型研究者阅读。
近期研究表明,仅用约1000条推理轨迹进行数据高效的后训练,也能在大语言模型中激发非平凡的推理能力。这类训练语料常包含'wait'、'so'、'alternatively'等标志性话语标记词,它们频繁出现在推理轨迹中,可能在其中起关键作用。本文聚焦于后训练中可观察的词汇级模式,通过案例对比数据高效的监督微调(SFT)与大规模后训练的差异。首先,我们识别出在不同模型和训练设置下,与正确答案相关的词汇模式。随后,重点研究'wait'标记词的分布及其功能角色,比较数据高效训练模型与大规模训练模型的表现。研究发现,话语标记词与推理准确性显著相关,并在小样本SFT中也带来准确率跃升。这表明小样本SFT可部分复现话语模式以模仿有意义的推理行为,但其模式与高置信度答案转换的对齐程度,仍低于大规模后训练的结果。
原文摘要 · Abstract (English)
Recent studies suggest that even data-efficient training with ($\simeq$1K) reasoning trajectories can induce non-trivial reasoning capabilities in large language models through post-training. Such training corpora often contain iconic tokens such as "wait", "so", and "alternatively", which frequently appear in reasoning trajectories and may play a role in this process. This paper focuses on characterizing observable token-level patterns in post-training and a case study of how data-efficient supervised fine-tuning (SFT) differs from, and falls short of, large-scale post-training. To this end, we first identify tokens that correlate with correct answers along reasoning trajectories across models and training setups. We then focus on the distribution and (functional) roles of the "wait" token to primarily study the model trained in a data-efficient manner compared with the counterpart. Our study finds that discourse tokens are associated with correctness and a reasoning accuracy jump, even in data-efficient SFT. This suggests data-efficient SFT can partially reproduce discourse-token patterns to mimic meaningful reasoning behavior, but the patterns are less aligned with high-confidence answer transitions than those from large-scale post-training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。