arXiv:2606.06491cs.ROcs.AI2026-06被引 2

让机器人在不同阶段自动变速,快慢由指令直接控制。

TempoVLA: Learning Speed-Controllable Vision-Language-Action Policies

论文配图:TempoVLA: Learning Speed-Controllable Vision-Language-Action Policies
图 1 · 摘自论文原文
  • 用动作重定时技术生成多速演示数据,保持动作语义不变
  • 模型根据速度指令实时调节执行快慢,仿真与实机均有效
  • 适合需要精细运动控制的机器人任务,如装配、抓取

机器人操作在低风险移动阶段需快速执行,高风险接触阶段则需缓慢精确。现有视觉-语言-动作模型(VLAs)仅继承单一固定速度,现有加速方法仅改变固定速度,几乎不支持减速。本文发现动作预测值本身即决定执行速度,由此提出TempoVLA——一个通过显式条件控制执行速度的单模型。其包含两个耦合组件:(1) 数据侧的可变速度轨迹增强(VSTA),通过合并或拆分动作重定时演示以匹配目标速度,保持动作语义;(2) 模型侧的速度条件机制。统计显示,VSTA能以极小运动误差达到目标速度。仿真与真实任务实验表明,TempoVLA实现双向灵活速度控制,且VSTA通过更好利用数据提升默认1×性能。结合大型多模态模型,可实现动态速度调节:低风险阶段加速,高风险阶段减速。

原文摘要 · Abstract (English)

Robot manipulation alternates between low-risk transit phases that call for fast execution and high-risk contact stages that demand slow, precise motion. Yet existing Vision-Language-Action models (VLAs) only inherit a single fixed speed from training demonstrations. Prior efforts to accelerate VLAs through model compression, KV-cache reuse, or reinforcement learning only shift the policy from one fixed speed to another, and leave deceleration almost unexplored. We observe that the magnitude of each predicted action already governs how fast the robot moves, opening a direct route to controllable execution speed. We turn this observation into TempoVLA, a single VLA whose execution speed is controlled by an explicit condition. TempoVLA combines two coupled components. (1) A data-side Variable-Speed Trajectory Augmentation (VSTA) that re-times demonstration to any target speed by merging or splitting actions while preserving its motion semantics. (2) A model-side conditioning mechanism that feeds the speed to the policy. Statistics show that VSTA reaches the requested speed with negligible motion error. Experiments in simulation and on real-world tasks demonstrate that TempoVLA achieves flexible speed control in both directions, while VSTA additionally boosts the default $1\times$ performance via better data utilization. Furthermore, by cooperating with a large multimodal model, TempoVLA realizes dynamic speed control, accelerating through low-risk phases and decelerating for high-risk ones.

机器人控制速度可控视觉语言动作动态调节

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。