arXiv:2607.12466cs.RO2026-07被引 1

用聚类方法从用户偏好中学习代表性奖励,提升机器人部署的适应性与稳定性。

Deployable Human Preference Alignment in Robotics: Learning Representative Rewards from Diverse Human Preferences

论文配图:Deployable Human Preference Alignment in Robotics: Learning Representative Rewards from Diverse Human Preferences
图 1 · 摘自论文原文
  • 基于用户偏好聚类,分组学习代表性奖励模型,减少个体差异干扰。
  • 在稀疏噪声反馈下,社会福利指标优于单一共享策略和逐用户对齐方法。
  • 适合需要兼顾多数与少数用户偏好的实际机器人部署场景。

将机器人策略与人类偏好对齐对于服务多样终端用户至关重要。单用户对齐方法因反馈稀疏导致学习不稳定且易受偏好噪声影响,同时需维护大量个性化策略,部署前验证困难。而单一共享策略虽降低复杂度,却无法捕捉偏好异质性,常忽略少数群体偏好。为此,本文提出基于偏好聚类的奖励学习框架(PREC),从多元用户提供的二元偏好标签中,学习一组紧凑的代表性策略。PREC首先剥离标签,跨用户聚合轨迹以训练全局共享的轨迹编码器,缓解单用户覆盖不足问题,并避免表示学习阶段受标签噪声影响。在此基础上,联合进行用户聚类与各簇内代表奖励模型学习,再为每簇优化对应策略。相似用户聚类可弥补单用户标签有限性并抑制噪声影响,同时保持少量奖励模型,降低部署验证负担。在多种模拟运动环境实验中,PREC比基线更准确地将不同轨迹标注者归入偏好一致的簇。在稀疏与噪声反馈下,其训练的策略在三项社会福利指标上均优于现有单一共享策略,甚至超越逐用户对齐方法。

原文摘要 · Abstract (English)

Aligning robot policies with human preferences is essential for deployment to diverse end users. In per-user alignment approach, preference feedback is often sparse, so learning becomes unstable and vulnerable to human preference noise, and a growing number of individualized policies makes validation difficult before deployment. A single shared policy approach to user alignment avoids this cost but fails to capture heterogeneous preferences and often neglects minority preferences. To address these challenges, we introduce Preference-based REward Clustering (PREC), a novel framework that learns a compact set of policies from binary preference labels provided by diverse users. From a dataset of user trajectories and their preference labels, PREC first sets the labels aside and aggregates trajectories across users to learn a population-level shared trajectory encoder, alleviating limited per-user coverage and avoiding label noise during representation learning. Using this representation, PREC jointly assigns users to preference-coherent clusters and learns a representative reward model per cluster using preference labels, from which a policy is optimized for each cluster. Clustering similar users compensates for the limited number of labels available from each user and mitigates the effect of label noise. At the same time, maintaining a manageable number of reward models reduces the validation burden at deployment. Experiments across diverse simulated locomotion environments show that PREC groups users who label different trajectory subsets into preference-coherent clusters more accurately than baseline methods. Under sparse and noisy feedback, policies trained with PREC improve all three social welfare metrics over an existing single shared-policy user-alignment approach and even outperform per-user alignment approaches.

机器人对齐偏好学习聚类奖励建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。