arXiv:2505.19767cs.RO2025-05被引 15

用时间反馈生成密集奖励,让机器人更聪明地完成复杂任务。

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback

  • 用时序信息训练价值模型,自动生成细粒度奖励。
  • 在CALVIN数据集上平均成功长度达4.296,创新高。
  • 仅需少量新环境训练,就能快速适应新任务。

视觉-语言-动作(VLA)模型在具身智能领域展现出巨大潜力,使智能体能在物理环境中遵循人类指令完成复杂任务。现有具身智能体通常通过行为克隆训练,需大量昂贵数据和计算资源,且受限于人类示范。为此,研究者尝试采用强化微调方法。然而,传统强化微调依赖稀疏的结果奖励,难以对回合内具体动作提供精细反馈,限制了模型的操作能力和泛化性能。本文提出RFTF,一种新型强化微调方法,利用价值模型生成具身场景中的密集奖励。该价值模型基于时间信息训练,无需昂贵的机器人动作标签。此外,RFTF融合GAE与样本平衡等技术,提升微调效果。实验表明,经RFTF微调的具身智能体在挑战性数据集CALVIN ABC-D上实现平均成功长度4.296的新纪录,并能快速适应新环境:在CALVIN D环境下仅经数次微调后,平均成功长度达4.301。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have demonstrated significant potential in the field of embodied intelligence, enabling agents to follow human instructions to complete complex tasks in physical environments. Existing embodied agents are often trained through behavior cloning, which requires expensive data and computational resources and is constrained by human demonstrations. To address this issue, many researchers explore the application of reinforcement fine-tuning to embodied agents. However, typical reinforcement fine-tuning methods for embodied agents usually rely on sparse, outcome-based rewards, which struggle to provide fine-grained feedback for specific actions within an episode, thus limiting the model's manipulation capabilities and generalization performance. In this paper, we propose RFTF, a novel reinforcement fine-tuning method that leverages a value model to generate dense rewards in embodied scenarios. Specifically, our value model is trained using temporal information, eliminating the need for costly robot action labels. In addition, RFTF incorporates a range of techniques, such as GAE and sample balance to enhance the effectiveness of the fine-tuning process. By addressing the sparse reward problem in reinforcement fine-tuning, our method significantly improves the performance of embodied agents, delivering superior generalization and adaptation capabilities across diverse embodied tasks. Experimental results show that embodied agents fine-tuned with RFTF achieve new state-of-the-art performance on the challenging CALVIN ABC-D with an average success length of 4.296. Moreover, RFTF enables rapid adaptation to new environments. After fine-tuning in the D environment of CALVIN for a few episodes, RFTF achieved an average success length of 4.301 in this new environment.

具身智能强化学习密集奖励快速适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。