arXiv:2602.12180cs.LGcs.GT2026-02被引 2

研究采样方式如何影响大模型对齐效果,揭示了动态迭代中的稳定性问题。

How Sampling Shapes LLM Alignment: From One-Shot Optima to Iterative Dynamics

  • 通过实例相关采样提升排序可靠性,避免过度集中。
  • 发现迭代对齐中存在持续振荡或熵崩溃现象。
  • 适用于关注大模型对齐机制与训练稳定性的研究者。

标准的大语言模型对齐方法基于采样候选响应的成对比较,并向参考策略正则化。尽管有效,采样与参考策略的选择在理论上仍不清晰。本文通过广泛使用的身份偏好优化(Identity Preference Optimization)框架进行研究,表明适当的实例依赖采样可获得更强的排序保证,而偏倚的在线采样可能在结构化偏好下引发过度集中。进一步分析了迭代对齐动态,其中学习到的策略反馈至后续采样与参考策略,反映了基于模型生成偏好数据的常见实践。理论证明,特定参数设置下该动态可能产生持续振荡或熵崩溃,并刻画了保证稳定性的区域。这些见解延伸至直接偏好优化(Direct Preference Optimization),表明所捕获的现象普遍存在于更广泛的偏好对齐方法中。真实世界偏好数据上的实验验证了上述发现。

原文摘要 · Abstract (English)

Standard methods for aligning large language models with human preferences learn from pairwise comparisons among sampled candidate responses and regularize toward a reference policy. Despite their effectiveness, the effects of sampling and reference choices are poorly understood theoretically. We investigate these effects through Identity Preference Optimization, a widely used preference alignment framework, and show that proper instance-dependent sampling can yield stronger ranking guarantees, while skewed on-policy sampling can induce excessive concentration under structured preferences. We then analyze iterative alignment dynamics in which the learned policy feeds back into future sampling and reference policies, reflecting a common practice of model-generated preference data. We prove that these dynamics can exhibit persistent oscillations or entropy collapse for certain parameter choices, and characterize regimes that guarantee stability. Our theoretical insights extend to Direct Preference Optimization, indicating the phenomena we captured are common to a broader class of preference-alignment methods. Experiments on real-world preference data validate our findings.

大模型对齐采样策略训练稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。