arXiv:2512.24673cs.ROcs.AI2025-12被引 12

解决视觉语言动作模型在机器人执行中的抖动与延迟问题,实现流畅高速控制。

VLA-RAIL: A Real-Time Asynchronous Inference Linker for VLA Models and Robots

  • 异步推理联动:模型与机器人运动解耦,提升响应效率。
  • 减少动作抖动,执行速度提升30%以上,任务成功率显著提高。
  • 适合需要连续高精度动作的机器人实操场景,如灵巧操作与动态任务。

视觉-语言-动作(VLA)模型在机器人领域取得显著进展,其中动作分块在这些突破中起主导作用。鉴于机器人运动控制具有实时性和连续性特征,连续动作分块的融合策略对整体性能影响深远。现有方法常导致动作执行出现抖动、卡顿甚至暂停,不仅限制了执行速度,也降低任务完成的成功率。本文提出VLA-RAIL(实时异步推理链接器),通过异步执行模型推理与机器人运动控制,保障动作执行的平滑、连续与高速。核心贡献包括:轨迹平滑器,利用多项式拟合有效过滤单个动作分块的噪声与抖动;分块融合器,无缝对齐当前执行轨迹与新到达的动作分块,确保位置、速度和加速度在相邻分块间的连续性。我们在动态仿真任务基准和多个真实世界操作任务上验证了VLA-RAIL的有效性。实验结果表明,该方法显著降低运动抖动,提升执行速度,并改善任务成功率,将成为VLA模型大规模部署的关键基础设施。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have achieved remarkable breakthroughs in robotics, with the action chunk playing a dominant role in these advances. Given the real-time and continuous nature of robotic motion control, the strategies for fusing a queue of successive action chunks have a profound impact on the overall performance of VLA models. Existing methods suffer from jitter, stalling, or even pauses in robotic action execution, which not only limits the achievable execution speed but also reduces the overall success rate of task completion. This paper introduces VLA-RAIL (A Real-Time Asynchronous Inference Linker), a novel framework designed to address these issues by conducting model inference and robot motion control asynchronously and guaranteeing smooth, continuous, and high-speed action execution. The core contributions of the paper are two fold: a Trajectory Smoother that effectively filters out the noise and jitter in the trajectory of one action chunk using polynomial fitting and a Chunk Fuser that seamlessly align the current executing trajectory and the newly arrived chunk, ensuring position, velocity, and acceleration continuity between two successive action chunks. We validate the effectiveness of VLA-RAIL on a benchmark of dynamic simulation tasks and several real-world manipulation tasks. Experimental results demonstrate that VLA-RAIL significantly reduces motion jitter, enhances execution speed, and improves task success rates, which will become a key infrastructure for the large-scale deployment of VLA models.

机器人控制异步推理动作生成轨迹优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。