arXiv:2505.10911cs.RO2025-05被引 63

无需新示范,语言指令就能让机器人学会新任务。

ReWiND: Language-Guided Rewards Teach Robot Policies without New Demonstrations

  • 用少量示范数据训练语言驱动的奖励模型,自动标注任务进展。
  • 在模拟和真实世界中,适应新任务的样本效率比基线高2~5倍。
  • 适合希望减少人工示范、实现快速部署的机器人研发人员。

我们提出 ReWiND 框架,仅通过语言指令即可学习机器人操作任务,无需每项任务都提供示范。传统强化学习与模仿学习需为每个新任务设计奖励函数或提供人类示范,而 ReWiND 从少量示范数据出发,学习:(1) 一种数据高效的、基于语言的奖励函数,用于对数据集打分;(2) 用该奖励函数预训练的语言条件策略。面对未见过的任务变体时,ReWiND 使用学习到的奖励函数微调预训练策略,仅需极少在线交互。实验表明,其奖励模型在未见任务上泛化能力优异,奖励泛化与策略对齐指标优于基线最高达2.4倍。此外,ReWiND 在模拟环境中使新任务适应效率提升2倍,在真实世界中使双臂预训练策略性能提升5倍,推动可扩展、面向现实世界的机器人学习发展。

原文摘要 · Abstract (English)

We introduce ReWiND, a framework for learning robot manipulation tasks solely from language instructions without per-task demonstrations. Standard reinforcement learning (RL) and imitation learning methods require expert supervision through human-designed reward functions or demonstrations for every new task. In contrast, ReWiND starts from a small demonstration dataset to learn: (1) a data-efficient, language-conditioned reward function that labels the dataset with rewards, and (2) a language-conditioned policy pre-trained with offline RL using these rewards. Given an unseen task variation, ReWiND fine-tunes the pre-trained policy using the learned reward function, requiring minimal online interaction. We show that ReWiND's reward model generalizes effectively to unseen tasks, outperforming baselines by up to 2.4x in reward generalization and policy alignment metrics. Finally, we demonstrate that ReWiND enables sample-efficient adaptation to new tasks, beating baselines by 2x in simulation and improving real-world pretrained bimanual policies by 5x, taking a step towards scalable, real-world robot learning. See website at https://rewind-reward.github.io/.

机器人学习语言引导零样本适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。