arXiv:2607.08877cs.ROcs.LG2026-07被引 1

用人类干预在潜在空间快速适应生成式机器人策略,无需大量数据或硬件训练。

FlowDAgger: Human-in-the-Loop Adaptation of Generative Robot Policies in Latent Space

论文配图:FlowDAgger: Human-in-the-Loop Adaptation of Generative Robot Policies in Latent Space
图 1 · 摘自论文原文
  • 通过逆向推导人类操作对应的噪声,实现对冻结模型的轻量级潜空间微调。
  • 仅需少量干预即在仿真与真实双臂/单臂任务中显著提升成功率,优于监督微调和强化学习基线。
  • 适合需要快速安全适配机器人基础模型的实际场景,保留预训练技能不退化。

基于流匹配和扩散模型的预训练生成式机器人策略在多种操作任务中表现优异,但实际部署常暴露于预训练分布外的失败模式。传统补救方法需大规模数据收集或在线强化学习,难以实现快速安全适配。本文提出FlowDAgger,一种在潜在空间中利用人类干预进行高效适应的方法。核心思想是动作逆推:通过反向时间积分结合局部优化,将人类专家动作映射为原冻结基模型生成该动作所对应的噪声。该逆推噪声作为监督信号,训练轻量级潜空间策略,在部署时引导基模型,实现快速技能习得并保留其行为先验。我们在仿真及真实世界双臂和单臂操作中评估了FlowDAgger,仅需少量干预即可适应基于动作头的VLAs和世界动作模型。结果表明,FlowDAgger优于监督微调与潜空间强化学习基线,并在未见任务上保持预训练技能,为真实世界中机器人基础模型的实用化适配提供了可行路径。

原文摘要 · Abstract (English)

Pretrained generative robot policies based on flow matching and diffusion have achieved impressive results across a wide range of manipulation tasks. Yet real-world deployments routinely expose failure modes outside the pretraining distribution. Closing these gaps typically requires large-scale data collection or online reinforcement learning on physical hardware, which is impractical for rapid and safe adaptation. We present FlowDAgger, a sample- and compute-efficient method for adapting frozen generative robot policies from human interventions in latent space. Our key idea is action inversion: each human expert action is mapped to the noise that would have produced it under the frozen base policy, using reverse-time integration followed by local refinement. The resulting inverted noise provides supervision for a lightweight latent policy that steers the base model at deployment time, enabling rapid skill acquisition while preserving its behavioral priors. We evaluate FlowDAgger in simulation and on real-world bimanual and single-arm manipulation, adapting both action-head VLAs and world-action models from a handful of interventions. FlowDAgger outperforms supervised fine-tuning and latent-space RL baselines and preserves pretrained skills on held-out tasks, offering a practical path for adapting robot foundation models in the real world. Website: https://microsoft.github.io/FlowDAgger

机器人生成模型人机协同潜在空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。