人类参与的双角色微调让机器人更聪明地完成复杂任务
Dual-Actor Fine-Tuning of VLA Models: A Talk-and-Tweak Human-in-the-Loop Approach
- 用两个角色协同:主角色保多任务能力,副角色学人类修正
- 真实场景下101分钟内三任务全成功,长序列任务成功率50%
- 适合需要人类指导的机器人学习场景,尤其多机器人协作
视觉-语言-动作(VLA)模型在机器人操作中展现出强大泛化能力,但在复杂现实任务中仍面临挑战。监督微调受限于数据质量,强化学习(RL)提供了可行替代方案。本文提出一种基于强化学习的人类在环双角色微调框架,融合主角色以保障多任务鲁棒性,以及优化角色用于隐空间适应。除物理干预外,引入轻量级‘对话修正’机制,将人类反馈转化为语义明确的语言指令,生成新策略学习数据集。真实世界多任务实验显示,仅需101分钟在线微调即实现三项任务100%成功率;长时序任务中维持连续12次操作50%成功率。该框架可有效扩展至多机器人训练,双机器人设置下效率最高提升2倍。实验视频见https://sites.google.com/view/hil-daft/
原文摘要 · Abstract (English)
Vision-language-action (VLA) models demonstrate strong generalization in robotic manipulation but face challenges in complex, real-world tasks. While supervised fine-tuning with demonstrations is constrained by data quality, reinforcement learning (RL) offers a promising alternative. We propose a human-in-the-loop dual-actor fine-tuning framework grounded in RL. The framework integrates a primary actor for robust multi-task performance with a refinement actor for latent-space adaptation. Beyond standard physical interventions, we introduce a lightweight talk-and-tweak scheme that converts human corrections into semantically grounded language commands, thereby generating a new dataset for policy learning. In real-world multi-task experiments, our approach achieves 100% success across three tasks within 101 minutes of online fine-tuning. For long-horizon tasks, it sustains a 50% success rate over 12 consecutive operations. Furthermore, the framework scales effectively to multi-robot training, achieving up to a 2 times improvement in efficiency when using dual robots. The experiment videos are available at https://sites.google.com/view/hil-daft/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。