arXiv:2603.09011cs.ROcs.AI2026-03

让机器人通过更人性化的排序任务学习用户偏好,提升交互体验。

Improving through Interaction: Searching Behavioral Representation Spaces with CMA-ES-IG

  • 用信息增益优化轨迹选择,使用户排序更直观
  • 在高维偏好空间中表现更好,且对噪声反馈鲁棒
  • 非专家用户更愿意使用,适合人机协作场景

在以人为中心的环境中,机器人需适应用户的个性化偏好才能有效运行。一种直观有效的学习非专家用户偏好的方法是通过行为(如轨迹、手势或语音)的排序。现有方法主要关注优化学习结果,如样本效率或最终估计精度,但忽略了用户在排序过程中的体验,可能影响系统采纳率。本文提出协方差矩阵自适应进化策略结合信息增益(CMA-ES-IG),在偏好学习中显式考虑用户体验,建议感知差异明显且信息丰富的轨迹供用户排序。通过模拟与真实机器人实验验证,CMA-ES-IG相比现有最优方法:(1) 更好地扩展至高维偏好空间;(2) 保持高维问题下的计算可行性;(3) 对噪声或不一致用户反馈具有鲁棒性;(4) 获得非专家用户更青睐,能更准确识别其偏好。代码已开源:github.com/interaction-lab/CMA-ES-IG。

原文摘要 · Abstract (English)

Robots that interact with humans must adapt to individual users' preferences to operate effectively in human-centered environments. An intuitive and effective technique to learn non-expert users' preferences is through rankings of robot behaviors, e.g., trajectories, gestures, or voices. Existing techniques primarily focus on generating queries that optimize preference learning outcomes, such as sample efficiency or final preference estimation accuracy. However, the focus on outcome overlooks key user expectations in the process of providing these rankings, which can negatively impact users' adoption of robotic systems. This work proposes the Covariance Matrix Adaptation Evolution Strategies with Information Gain (CMA-ES-IG) algorithm. CMA-ES-IG explicitly incorporates user experience considerations into the preference learning process by suggesting perceptually distinct and informative trajectories for users to rank. We demonstrate these benefits through both simulated studies and real-robot experiments. CMA-ES-IG, compared to state-of-the-art alternatives, (1) scales more effectively to higher-dimensional preference spaces, (2) maintains computational tractability for high-dimensional problems, (3) is robust to noisy or inconsistent user feedback, and (4) is preferred by non-expert users in identifying their preferred robot behaviors. This project's code is available at github.com/interaction-lab/CMA-ES-IG

人机交互偏好学习进化算法机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。