arXiv:2603.01581cs.ROcs.LG2026-03中稿 · DAC 2026被引 8

KERV通过融合机器人运动学提升视觉语言动作模型推理速度,无需重推理。

KERV: Kinematic-Rectified Speculative Decoding for Embodied VLA Models

  • 用运动学卡尔曼滤波预测动作并修正生成误差
  • 动态调整接受阈值,实现27%~37%加速且成功率几乎不变
  • 适合需实时控制的具身智能任务,如机器人操作

视觉-语言-动作(VLA)模型构建了基于标记域的机器人控制范式,但存在推理速度低的问题。推测解码(SD)是一种可提升推理速度的优化策略。然而将VLA与SD结合时出现两大问题:一是SD依赖重推理来修正标记错误,计算开销大;二是为降低标记错误,接收阈值需精细调节,现有方法未能有效解决。同时,作为人工智能与物理世界的桥梁,现有具身智能忽略了机器人运动学的应用。为此,我们创新性地将标记域的VLA模型与运动学域预测相结合,提出名为KERV的运动学修正推测解码框架。采用基于运动学的卡尔曼滤波器预测动作并补偿解码误差,避免高成本重推理;同时设计基于运动学的阈值动态调整策略,缓解阈值设定难题。跨多种任务与环境的实验结果表明,KERV在保持成功率几乎不变的前提下,实现了27%~37%的加速。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models build a token-domain robot control paradigm, yet suffer from low speed. Speculative Decoding (SD) is an optimization strategy that can boost inference speed. Two key issues emerge when integrating VLA and SD: first, SD relies on re-inference to address token errors, which is computationally expensive; second, to mitigate token errors, the acceptance threshold in SD requires careful adjustment. Existing works fail to address the above two issues effectively. Meanwhile, as the bridge between AI and the physical world, existing embodied intelligence has overlooked the application of robotic kinematics. To address these issues, we innovatively combine token-domain VLA models with kinematic-domain prediction for SD, proposing a kinematic-rectified SD framework named KERV. We employ a kinematics-based Kalman Filter to predict actions and compensate for SD errors, avoiding costly re-inference. Moreover, we design a kinematics-based adjustment strategy to dynamically rectify the acceptance threshold, addressing the difficulty of threshold determination. Experimental results across diverse tasks and environments demonstrate that KERV achieves 27%~37% acceleration with nearly no Success Rate loss.

具身智能推测解码运动学机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。