arXiv:2602.20457cs.LGstat.ML2026-02被引 1

让大模型在错误反馈下仍能稳定对齐,提升训练鲁棒性。

Oracle-Robust Online Alignment for Large Language Models

  • 构建反馈不确定集,以最坏情况优化对齐目标。
  • 理论证明可分解为原损失+敏感度惩罚项。
  • 适合追求训练稳定性的大模型研究者使用。

我们研究在偏好反馈不准确情形下的大语言模型在线对齐问题,其中观测到的偏好代理与未知的理想真实代理存在偏差。由于数据收集与策略更新的耦合,该问题本质上是双层强化学习问题。近期,SAIL(自改进高效在线对齐)框架将其简化为可处理的单层目标。本文引入点态偏好不确定性集,将对齐目标建模为最坏情况优化问题。对于对数线性策略,我们证明该鲁棒目标可精确分解为原始损失函数加上显式敏感度惩罚项。我们设计了投影随机复合更新算法来处理所得弱凸目标,并证明达到近似驻点的Oracle复杂度为$\widetilde{O}(\varepsilon^{-2})$。

原文摘要 · Abstract (English)

We study online alignment of large language models under misspecified preference feedback, where the observed preference oracle deviates from an ideal but unknown ground-truth oracle. The online LLM alignment problem is a bi-level reinforcement problem due to the coupling between data collection and policy updates. Recently, the problem has been reduced to tractable single-level objective in the SAIL (Self-Improving Efficient Online Alignment) framework. In this paper, we introduce a pointwise oracle uncertainty set in this problem and formulate an oracle-robust online alignment objective as a worst-case optimization problem. For log-linear policies, we show that this robust objective admits an exact closed-form decomposition into the original loss function plus an explicit sensitivity penalty. We develop projected stochastic composite updates for the resulting weakly convex objective and prove $\widetilde{O}(\varepsilon^{-2})$ oracle complexity for reaching approximate stationarity.

大模型对齐在线学习鲁棒优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。