用视觉示范+语言指导,让机器人更好理解用户意图
ELEMENTAL: Interactive Learning from Demonstrations and Vision-Language Models for Reward Design in Robotics
- 结合视觉示范与自然语言,动态调整奖励权重
- 任务成功率提升42.3%,分布外任务泛化能力提高41.3%
- 适合非专家用户快速设计机器人奖励函数
强化学习在机器人任务中表现优异,但其成功常依赖复杂且人为设计的奖励函数。尽管大型语言模型(LLMs)可帮助非专家更易指定奖励函数,但仍存在难以平衡特征重要性、对分布外任务泛化能力差、仅靠文本描述无法准确表征问题等缺陷。为此,我们提出ELEMENTAL(intEractive LEarning froM dEmoNstraTion And Language)框架,融合自然语言指导与视觉用户示范,更好地对齐机器人行为与用户意图。通过引入视觉输入,克服纯文本描述的局限性;利用逆强化学习(IRL)优化特征权重并匹配示范行为。此外,引入自省式迭代反馈循环,持续改进特征、奖励与策略学习。实验表明,ELEMENTAL在任务成功率上比现有方法提升42.3%,在分布外任务上的泛化性能提升41.3%,展现出更强的演示学习鲁棒性。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has demonstrated compelling performance in robotic tasks, but its success often hinges on the design of complex, ad hoc reward functions. Researchers have explored how Large Language Models (LLMs) could enable non-expert users to specify reward functions more easily. However, LLMs struggle to balance the importance of different features, generalize poorly to out-of-distribution robotic tasks, and cannot represent the problem properly with only text-based descriptions. To address these challenges, we propose ELEMENTAL (intEractive LEarning froM dEmoNstraTion And Language), a novel framework that combines natural language guidance with visual user demonstrations to align robot behavior with user intentions better. By incorporating visual inputs, ELEMENTAL overcomes the limitations of text-only task specifications, while leveraging inverse reinforcement learning (IRL) to balance feature weights and match the demonstrated behaviors optimally. ELEMENTAL also introduces an iterative feedback-loop through self-reflection to improve feature, reward, and policy learning. Our experiment results demonstrate that ELEMENTAL outperforms prior work by 42.3% on task success, and achieves 41.3% better generalization in out-of-distribution tasks, highlighting its robustness in LfD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。