提升机器人动作连贯性,让视觉语言指令更稳定可靠
ACG: Action Coherence Guidance for Flow-based Vision-Language-Action models
- 测试时无需训练,通过引导增强动作连贯性
- 在多个真实任务中成功率显著提升,动作更平滑
- 适合需要精细操作的机器人部署场景
扩散模型和流匹配模型已成为强大的机器人策略,使视觉-语言-动作(VLA)模型能够在多样场景和指令下泛化。然而,通过模仿学习训练时,其高生成能力使其对人类示范中的噪声敏感:如动作抖动、停顿和抖动,导致动作连贯性下降。连贯性降低会引发部署时的不稳定性与轨迹漂移,在精细操作任务中可能造成灾难性失败。本文提出动作连贯性引导(Action Coherence Guidance, ACG),一种无需训练的测试时引导算法,可有效提升动作连贯性并带来性能提升。在RoboCasa、DexMimicGen和真实世界SO-101任务上评估,ACG在多种操纵任务中一致提升动作连贯性和成功率。代码与项目页见https://github.com/DAVIAN-Robotics/ACG 和 https://DAVIAN-Robotics.github.io/ACG。
原文摘要 · Abstract (English)
Diffusion and flow matching models have emerged as powerful robot policies, enabling Vision-Language-Action (VLA) models to generalize across diverse scenes and instructions. Yet, when trained via imitation learning, their high generative capacity makes them sensitive to noise in human demonstrations: jerks, pauses, and jitter which reduce action coherence. Reduced action coherence causes instability and trajectory drift during deployment, failures that are catastrophic in fine-grained manipulation where precision is crucial. In this paper, we present Action Coherence Guidance (ACG) for VLA models, a training-free test-time guidance algorithm that improves action coherence and thereby yields performance gains. Evaluated on RoboCasa, DexMimicGen, and real-world SO-101 tasks, ACG consistently improves action coherence and boosts success rates across diverse manipulation tasks. Code and project page are available at https://github.com/DAVIAN-Robotics/ACG and https://DAVIAN-Robotics.github.io/ACG , respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。