arXiv:2601.01624cs.CL2026-01

prefix能提升推理模型的安全性和逻辑性,效果优于传统方法。

How Does Prefix Matter in Reasoning Model Tuning?

  • 用带前缀的指令微调模型,引导其生成更安全、连贯的回答。
  • 在对抗性测试中安全率提升6%,数学推理准确率提高7%。
  • 适合关注模型安全与逻辑推理优化的研究者和工程师。

近期对齐研究通常从监督微调(SFT)数据集中移除开头的通用语句。本文挑战这一假设,提出安全与推理导向的前缀句可作为轻量级对齐信号,引导模型解码生成更安全、更连贯的响应。我们对三个R1系列模型在三大核心能力(数学、编码、安全、事实性)上系统性地调整前缀比例(0%至100%)。结果表明,带前缀的SFT显著提升安全与推理性能,在对抗性基准(WildJailbreak、StrongReject)上安全率最高提升6%,在GSM8K推理任务上提升7%。但事实性与编码任务表现趋弱或下降,说明前缀缩小搜索空间有利于结构化推理。词元级损失分析显示,'revised'、'logically'等前缀词具有更高梯度幅值,充当对齐锚点,稳定推理轨迹。研究证明,前缀条件化是一种可扩展、可解释的对齐机制,可补充传统基于奖励的方法。

原文摘要 · Abstract (English)

Recent alignment studies commonly remove introductory boilerplate phrases from supervised fine-tuning (SFT) datasets. This work challenges that assumption. We hypothesize that safety- and reasoning-oriented prefix sentences serve as lightweight alignment signals that can guide model decoding toward safer and more coherent responses. To examine this, we fine-tune three R1 series models across three core model capabilities: reasoning (mathematics, coding), safety, and factuality, systematically varying prefix inclusion from 0% to 100%. Results show that prefix-conditioned SFT improves both safety and reasoning performance, yielding up to +6% higher Safe@1 accuracy on adversarial benchmarks (WildJailbreak, StrongReject) and +7% improvement on GSM8K reasoning. However, factuality and coding tasks show marginal or negative effects, indicating that prefix-induced narrowing of the search space benefits structured reasoning. Token-level loss analysis further reveals that prefix tokens such as "revised" and "logically" incur higher gradient magnitudes, acting as alignment anchors that stabilize reasoning trajectories. Our findings suggest that prefix conditioning offers a scalable and interpretable mechanism for improving reasoning safety, serving as an implicit form of alignment that complements traditional reward-based methods.

模型对齐推理增强安全生成微调策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。