让视觉语言动作模型支持动态追踪,提升机器人操作速度与安全性
Enabling Dynamic Tracking in Vision-Language-Action Models via Time-Discrete and Time-Continuous Velocity Feedforward
- 通过离散差分或连续样条方式提取速度指令,融入现有模型
- 在立方体插孔任务中,速度前馈使执行速度显著提升,成功率保持高位
- 无需修改模型结构,适配工业机器人高频率控制需求
尽管视觉语言动作(VLA)模型在机器人操作中展现出巨大潜力,但其在刚性工业机器人上的部署仍面临顺应性与响应速度之间的权衡难题。标准行为克隆(BC)方法以低频预测离散位姿,忽略了低层柔顺控制器常用的速度与加速度前馈项,导致必须依赖高刚度实现精确跟踪,牺牲了安全接触特性。本文证明,在VLA策略中集成速度前馈项可有效解决该矛盾。提出两种从VLA中提取速度目标的方法:一种是时间离散的有限差分近似,可作为现有模型的有效桥梁;另一种是连续三次样条动作空间,原生生成$C^2$连续轨迹,适用于高频控制。关键的是,两种方法均完全模型无关,兼容任何标准动作分块架构,仅需调整遥操作、数据处理及底层控制器。我们对$π_{0.5}$模型进行微调,并在高接触密度的立方体插孔任务上评估两种方法。结果表明,通过有限差分引入速度前馈能显著提升任务执行速度;而样条方法在保持高成功率的同时,为更高阶导数平滑提供基础,且不牺牲顺应性。
原文摘要 · Abstract (English)
While vision-language-action (VLA) models have shown great promise for robot manipulation, their deployment on rigid industrial robots remains challenging due to the inherent trade-off between compliance and responsiveness. Standard Behavior Cloning (BC) approaches predict discrete poses at low frequencies, omitting the velocity and acceleration feedforward terms typically used by low-level compliant controllers. This requires to rely on high stiffness for accurate tracking, thereby sacrificing safe contact dynamics. In this paper, we demonstrate the importance of integrating velocity feedforward terms into VLA policies to resolve this trade-off. We propose two methods for extracting velocity targets from VLAs: a time-discrete finite-difference approximation that serves as a highly effective bridge for existing models, and a continuous Cubic B-Spline action space that natively yields $C^2$ continuous trajectories for high-frequency control. Crucially, both approaches are strictly model-agnostic and compatible with any standard action-chunking architecture, requiring modifications only to teleoperation, data processing, and the low-level controller. We fine-tune the $π_{0.5}$ model and evaluate both of our approaches on a demanding, contact-rich cube-in-hole task. Our results indicate that incorporating the velocity feedforward term via finite differences significantly improves task execution speed, while the continuous B-Spline approach maintains high overall success rates and provides a foundation for smoother higher-order derivatives without compromising compliance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。