arXiv:2608.18770cs.LGcs.RO2026-08

通过构建用户偏好课程,提升个性化对齐中难优化用户的满意度。

To Go Far, Go Together: Diverse Preferences Induce a Curriculum for Reward Optimization

论文配图:To Go Far, Go Together: Diverse Preferences Induce a Curriculum for Reward Optimization
图 1 · 摘自论文原文
  • 基于用户偏好难易度构建树状课程,动态适应不同用户需求。
  • 在模拟环境中使群体满意度提升1.2至2.1倍,训练时间显著缩短。
  • 适合关注个性化对齐与公平性的研究人员,尤其针对少数群体。

从公平性角度出发,现有方法通过高效准确的奖励模型捕捉少数用户偏好。本文进一步指出:仅建模准确且数据高效还不够,若用户奖励模型在策略优化阶段难以被优化,仍会成为新弱势群体。观察发现,部分用户奖励模型从初始策略即可轻松优化,而另一些则困难。在足够多样化的用户群体中,自然形成从易到难的优化课程。为此提出CurriPO,构建树状课程结构,支持分支扩展并复用已有奖励模型,实现单次遍历覆盖全人群。实验表明,在模拟个性化连续控制任务中,该方法使群体满意度达到最强基线的1.2–2.1倍,同时大幅减少训练时间。分析显示,主要提升源于传统方法忽视的未被充分服务用户。

原文摘要 · Abstract (English)

Learning a reward model from human feedback and optimizing a policy against it is one approach to aligning AI systems with individual users. From a fairness perspective, existing work improves such alignment by developing data-efficient and accurate reward models that capture minority preferences despite scarce data. We push this line of inquiry one step further and argue that data-efficient and accurate per-user reward models are not sufficient: users whose reward models are difficult to \textit{optimize} at the policy level can become a new underserved group. We start from the observation that one user's reward model can be easy to optimize from the initial policy while another's is not. We argue that, given a sufficiently diverse user population, a curriculum naturally emerges between easy- and hard-to-optimize reward models. Building on this insight, we propose CurriPO, which grows a tree-structured curriculum to accommodate diverse user-specific objectives, covering the population in a single traversal. Specifically, CurriPO automatically constructs a curriculum over diverse user reward models, allowing it to branch from the existing curriculum and reuse reward models previously incorporated into the curriculum. To the best of our knowledge, this is the first work to explicitly exploit multi-user structure to address optimization in AI alignment. Extensive experiments on personalized continuous control in a simulated environment show that CurriPO achieves $1.2$--$2.1\times$ the population satisfaction of the strongest baseline while substantially reducing training time. Additional analysis attributes much of this improvement to the users left underserved by conventional optimization.

AI对齐个性化课程学习奖励建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。