arXiv:2607.07023cs.LG2026-07

在线数据选择会悄悄改变模型行为,无需强化学习也能实现对齐。

Online Data Selection Is Implicit Alignment

论文配图:Online Data Selection Is Implicit Alignment
图 1 · 摘自论文原文
  • 将训练数据选择过程视为隐式对齐机制,用评分权重替代奖励模型。
  • 不同选择策略导致拒绝率、冗余度和顺从性显著差异,但任务准确率相近。
  • 适合关注模型安全与风格偏移的研究者,尤其在数据高效训练场景中。

监督微调(SFT)常被视为能力适配步骤,而对齐则归因于后续的偏好优化或强化学习。这种划分并不完整:当在微调过程中在线评分并筛选数据时,选择哪些数据训练已改变了模型的行为偏好。本文研究在线数据选择作为隐式对齐机制的作用。在相同基模型、优化器和选中令牌预算下,对比随机、基于损失、基于质量及基于多样性的在线选择器,并测量其在无偏好优化情况下引发的行为漂移。评估涵盖帮助性、拒绝率、冗余度、真实性、奉承倾向、校准性及越狱鲁棒性,同时诊断所选数据中何种行为模式占主导。我们形式化在线选择为重加权SFT目标,其权重定义了响应风格与安全姿态的隐式偏好,使在线评分器承担原本由奖励模型扮演的角色。该视角预测:高分数据可能系统性偏向更长、更自信、更顺从或更易拒绝的行为,具体取决于评分定义。实证表明,统计上不可区分的任务准确率下,各选择器在拒绝率、冗余度和奉承性上表现迥异,且漂移方向可由所选数据的属性混合预测。我们提出对齐漂移审计(ADA),一种量化选择诱导行为变化的受控协议;以及对齐感知选择(AAS),一种在保持数据效率的同时约束安全与风格漂移的诊断型选择器。

原文摘要 · Abstract (English)

Supervised fine-tuning (SFT) is often treated as a capability-adaptation step, while alignment is attributed to later preference optimization or reinforcement learning. This separation is incomplete: when examples are scored and kept online during fine-tuning, the choice of which data to train on already changes the model's behavioral preferences. We study online data selection as an implicit alignment mechanism. Given the same base model, optimizer, and selected-token budget, we compare random, loss-based, quality-based, and diversity-based online selectors and measure the behavioral drift they induce without any preference optimization. The proposed evaluation tracks helpfulness, refusal rate, verbosity, truthfulness, sycophancy, calibration, and jailbreak robustness, together with diagnostics for which behavioral modes are over-represented in the selected data. We formalize online selection as a reweighted SFT objective whose weights define an implicit preference over response styles and safety postures, so that an online scorer plays the role usually assigned to a reward model. This view predicts that high-scoring data can systematically favor longer, more assertive, more compliant, or more refusal-prone behaviors depending on how the online score is defined. Empirically, selectors that are statistically indistinguishable in task accuracy diverge sharply in refusal rate, verbosity, and sycophancy, and we show that the direction of the shift is predictable from the attribute mixture of the selected data. We introduce Alignment Drift Auditing (ADA), a controlled protocol for quantifying selection-induced behavioral movement, and Alignment-Aware Selection (AAS), a diagnostic online selector that retains data efficiency while constraining drift along safety and style axes.

模型对齐数据选择行为漂移SFT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。