arXiv:2409.13822cs.RO2024-09被引 8

通过学习用户偏好动作表示,实现机器人高效个性化交互。

Personalization in Human-Robot Interaction through Preference-based Action Representation Learning

  • 用预训练模型做参考,学习用户偏好的动作表征
  • 在8人真实实验中,个性化效果优于现有方法
  • 无需重训,保留原任务性能,适合实际应用

基于偏好的强化学习(PbRL)在人机交互个性化中表现优异,能将人类偏好显式融入机器人学习过程。但现有方法通常需从头训练个性化策略,导致人类反馈利用效率低。本文提出偏好驱动的动作表征学习(PbARL),一种高效微调方法,通过解耦通用任务结构与用户偏好,利用预训练机器人策略作为参考。PbARL不直接微调策略,而是将其用于最大化源域与目标偏好对齐域之间的互信息动作表征学习。该方法使机器人在保持原有任务性能的同时实现个性化行为调整,且无需源域的大量先验信息,显著提升实际人机交互场景中的效率与实用性。在Assistive Gym基准和一项包含8名用户的实地研究中,结果表明本方法优于当前最优方法。

原文摘要 · Abstract (English)

Preference-based reinforcement learning (PbRL) has shown significant promise for personalization in human-robot interaction (HRI) by explicitly integrating human preferences into the robot learning process. However, existing practices often require training a personalized robot policy from scratch, resulting in inefficient use of human feedback. In this paper, we propose preference-based action representation learning (PbARL), an efficient fine-tuning method that decouples common task structure from preference by leveraging pre-trained robot policies. Instead of directly fine-tuning the pre-trained policy with human preference, PbARL uses it as a reference for an action representation learning task that maximizes the mutual information between the pre-trained source domain and the target user preference-aligned domain. This approach allows the robot to personalize its behaviors while preserving original task performance and eliminates the need for extensive prior information from the source domain, thereby enhancing efficiency and practicality in real-world HRI scenarios. Empirical results on the Assistive Gym benchmark and a real-world user study (N=8) demonstrate the benefits of our method compared to state-of-the-art approaches.

人机交互个性化强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。