arXiv:2601.17260cs.LGcs.AI2026-01

发现DPO对齐中存在逻辑能力突变与滞后现象,高对齐压力可能反而损害模型能力。

The Viscosity of Logic: Phase Transitions and Hysteresis in DPO Alignment

  • 将对齐强度β作为控制变量,系统扫描其影响
  • β≈10⁻²时逻辑能力达峰值,过高或过低均导致能力下降
  • 首次揭示对齐边际与推理能力负相关,适合关注模型可靠性的研究者

Direct Preference Optimization (DPO) 通常被当作提升对齐强度(由β控制)会持续改善行为。本文将β视为控制参数,在固定DPO流程下,对三个7B规模的开源模型家族进行密集扫描。在Mistral中,能力呈现尖锐非单调变化:仅在β≈10⁻²附近,逻辑探测得分呈正向增长,且边界点对随机种子敏感。不同架构响应模式各异:Mistral出现剧烈重构,Llama为选择性调整,Qwen则呈现平滑权衡。关键发现:DPO偏好边际与推理能力呈强负相关(皮尔逊r = -0.91,Llama逻辑任务),意味着基于边际的选择可能偏好能力受损模型。训练路径亦有影响:曾暴露于高β的模型即使降低β后,能力损失仍持续存在(滞后效应)。这些结果强调应跨β空间开展能力分辨评估,而非依赖边际或整体基准。

原文摘要 · Abstract (English)

Direct Preference Optimization (DPO) is often tuned as if increasing alignment pressure (controlled by $β$) yields progressively "better" behavior. We instead treat $β$ as a control parameter and densely sweep it for three 7B open-weight families under a fixed DPO recipe. In Mistral, capability is sharply non-monotonic: aggregated logic-probe margins become positive only in a narrow band near $β\approx 10^{-2}$ and revert outside it, with boundary points that are seed-sensitive. Across architectures under the same sweep, we observe qualitatively different response modes: sharp reorganization in Mistral, selective changes in Llama, and smooth trade-offs in Qwen. Critically, the DPO preference margin can anticorrelate with reasoning capability (Pearson $r=-0.91$ for Llama logic), so margin-based selection can prefer capability-impaired models. Training path also matters: exposure to high $β$ induces capability losses that persist even after $β$ is reduced (hysteresis). These findings motivate capability-resolved evaluation across the $β$ landscape rather than reliance on margins or aggregate benchmarks.

对齐模型能力滞后效应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。