通过反向流动引导,让机器人通用策略更好理解人类指令并自动优化动作。
Improving Robotic Generalist Policies via Flow Reversal Steering

- 利用反向流模型从不合理动作推回潜在噪声,找到相近的优质动作模式。
- 结合视觉语言模型指令,零样本控制成功率提升最高达95%。
- 适合需要快速适应新任务的机器人应用,尤其擅长复杂场景下的动作优化。
通用策略可从多样化的机器人数据集中学习广泛技能。为解决或改进具有挑战性的新任务,需一种方法从策略丰富的行为先验中推断并调用合适动作,尤其当直接命令失败时。本文聚焦于流匹配型通用策略,提出流反转引导(FRS):该方法接收次优但合理的动作,通过反向传递至流策略以求解其潜在噪声,并将其映射到邻近的通用策略动作模式。我们在多个模拟与真实世界操作场景中评估了FRS。结果表明,FRS能将来自人类或视觉-语言模型(VLMs)的粗略语义指令转化为对应的良好机器人动作,显著提升零样本控制性能。这些改进可通过行为克隆进行提炼——训练一个辅助策略输出使通用策略生成优质动作的噪声,仅需一分钟训练即实现高达95%的绝对任务成功率提升。此外,FRS还能通过引入语义知识进行强化学习自举,改善标准强化学习无法提升的多个任务。
原文摘要 · Abstract (English)
Generalist policies can learn a wide range of skills from diverse robot datasets. In order to solve or improve on challenging new tasks, we need a way to infer and invoke the appropriate actions from the policy's rich behavioral prior, especially when directly commanding the policy fails. We focus on flow matching generalists and propose Flow Reversal Steering (FRS): a method that takes suboptimal but ``reasonable'' actions, finds their latent noises by passing them through the flow policy in reverse, and maps them to nearby generalist action modes. We evaluate FRS across many simulated and real-world manipulation settings. First, FRS can turn coarse semantic guidance from humans or vision-language models (VLMs) into corresponding good robot actions, improving zero-shot control. These gains can be distilled with behavioral cloning by training an auxiliary policy to output noises that the generalist maps to good actions -- showing up to 95% absolute task success rate boosts in under a minute of training. Finally, FRS enables policy improvement by bootstrapping reinforcement learning with semantic knowledge, improving on several tasks that standard RL fails to improve on.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。