arXiv:2605.20696cs.LG2026-05

解决分布式环境下偏好优化的收敛难题,提升模型对异构用户数据的适应性。

Distributed Direct Preference Optimization

  • 构建个性化离线强化学习框架,分析分布式场景下的全局优化结构。
  • 首次给出联邦与去中心化DPO的收敛速率,揭示通信频率与用户偏好的影响。
  • 理论结合实证,适合研究分布式强化学习与对齐算法的开发者参考。

基于偏好的强化学习是使策略与人类判断对齐的关键范式,但在偏好数据分散于异构用户的分布式设置中,其理论行为仍不明确。直接偏好优化(DPO)避免了显式的奖励建模,但在联邦和去中心化训练中缺乏收敛保证,因通信约束与非独立同分布偏好会从根本上改变优化动态。本文首次提供了DPO在分布式环境中的收敛性与时间复杂度分析。通过建模用户特定偏好分布的个性化离线强化学习,我们刻画了诱导出的全局优化景观。对于联邦DPO,推导出量化客户端漂移、通信频率和偏好异质性影响的收敛速率;对于去中心化DPO,建立了在一般通信图上的收敛性,并表明谱连通性决定优化速度与共识能力。实验上,在标准对齐基准上验证了理论洞察,证明所提方法不仅具备强理论保障,且在实践中表现鲁棒可扩展。代码已公开。

原文摘要 · Abstract (English)

Preference-based reinforcement learning (RL) is a key paradigm for aligning policies with human judgments, yet its theoretical behavior in distributed settings where preference data are fragmented across heterogeneous users remains poorly understood. Direct Preference Optimization (DPO) avoids explicit reward modeling but lacks convergence guarantees under federated and decentralized training, where communication constraints and non-IID preferences fundamentally alter optimization dynamics. We provide the first convergence and time-complexity analysis of DPO in distributed environments. Modeling personalized offline RL with user-specific preference distributions, we characterize the induced global optimization landscape. For federated DPO, we derive convergence rates that quantify the impact of client drift, communication frequency, and preference heterogeneity; for decentralized DPO, we establish convergence over general communication graphs and show how spectral connectivity governs optimization speed and consensus. Empirically, we corroborate our theoretical insights on standard alignment benchmarks, demonstrating that our proposed methods not only enjoy strong theoretical guarantees but also deliver robust and scalable performance in practice. The code base is available here.

强化学习分布式训练偏好优化联邦学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。