让模型同时满足不同用户偏好,发现平均奖励最有效
Direct Alignment with Heterogeneous Preferences
- 用用户类型建模人类偏好的多样性,以平均奖励实现统一对齐
- 只需少量标注信息即可显著提升对齐效果,全量反馈则能稳定学习最优策略
- 揭示了对齐中一致性和样本效率的根本矛盾,适合研究人机交互的学者
人类偏好本质上具有异质性,但现有对齐方法通常假设存在一个通用奖励函数。本文通过引入用户类型来形式化这种异质性,分析同质性假设的局限性。研究发现,使用单一策略对异质偏好进行对齐时,最优方式是采用跨用户类型的平均奖励。这需要额外的标注者信息。在不同信息条件下评估直接对齐方法,结果表明:最少信息即可带来一阶改进;而来自每类用户的完整反馈可确保最优策略的稳定学习。然而令人意外的是,在后一种情形下,不存在样本高效的稳定直接损失函数。这一发现揭示了直接策略对齐中一致性与样本效率之间的根本权衡。
原文摘要 · Abstract (English)
Alignment with human preferences is commonly framed using a universal reward function, even though human preferences are inherently heterogeneous. We formalize this heterogeneity by introducing user types and examine the limits of the homogeneity assumption. We show that aligning to heterogeneous preferences with a single policy is best achieved using the average reward across user types. However, this requires additional information about annotators. We examine improvements under different information settings, focusing on direct alignment methods. We find that minimal information can yield first-order improvements, while full feedback from each user type leads to consistent learning of the optimal policy. Surprisingly, however, no sample-efficient consistent direct loss exists in this latter setting. These results reveal a fundamental tension between consistency and sample efficiency in direct policy alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。