用对比性语言反馈优化机器人轨迹并学习人类偏好奖励函数
Trajectory Improvement and Reward Learning from Comparative Language Feedback
- 构建轨迹与语言反馈的共享潜在空间,融合多模态信息
- 实验显示方法主观评分高23.9%,效率提升11.3%
- 首次将对比语言反馈用于奖励学习,适合人机交互研究者
近年来,基于人类反馈的学习在机器人和自然语言处理领域日益受到关注。以往工作主要依赖比较形式的反馈,而语言能提供更丰富的用户偏好信息。本文旨在利用对比性语言反馈,迭代改进机器人轨迹并学习编码人类偏好的奖励函数。为此,我们学习一个整合轨迹数据与语言反馈的共享潜在空间,并基于该空间优化轨迹与学习偏好。据我们所知,这是首个将对比语言反馈引入奖励学习的工作。仿真实验验证了潜在空间的有效性及算法的成功。人机实验表明,本方法平均主观评分高出23.9%,时间效率提升11.3%,显著优于基于偏好的奖励学习。相关网站见 https://liralab.usc.edu/comparative-language-feedback/
原文摘要 · Abstract (English)
Learning from human feedback has gained traction in fields like robotics and natural language processing in recent years. While prior works mostly rely on human feedback in the form of comparisons, language is a preferable modality that provides more informative insights into user preferences. In this work, we aim to incorporate comparative language feedback to iteratively improve robot trajectories and to learn reward functions that encode human preferences. To achieve this goal, we learn a shared latent space that integrates trajectory data and language feedback, and subsequently leverage the learned latent space to improve trajectories and learn human preferences. To the best of our knowledge, we are the first to incorporate comparative language feedback into reward learning. Our simulation experiments demonstrate the effectiveness of the learned latent space and the success of our learning algorithms. We also conduct human subject studies that show our reward learning algorithm achieves a 23.9% higher subjective score on average and is 11.3% more time-efficient compared to preference-based reward learning, underscoring the superior performance of our method. Our website is at https://liralab.usc.edu/comparative-language-feedback/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。