arXiv:2606.14255cs.RO2026-06被引 1

让机器人快速反应:用单步生成动作,速度提升4倍以上

ReactVLA: Fast and Lightweight Reactive Robot Manipulation via Improved Mean Flow Action Generation

论文配图:ReactVLA: Fast and Lightweight Reactive Robot Manipulation via Improved Mean Flow Action Generation
图 1 · 摘自论文原文
  • 用改进的均值流方法,将多步采样压缩为1-3步动作生成
  • 在真实机器人上实现低于38.6毫秒延迟,精度任务性能提升1.65倍
  • 适合需要快速响应的工业机械臂、服务机器人等场景

基于扩散模型的视觉语言动作(VLA)策略在建模丰富多模态动作分布方面表现优异,但其依赖迭代采样导致推理延迟高,难以用于实时闭环机器人操作。为此,我们提出轻量级低延迟的VLA框架ReactVLA。ReactVLA结合两项互补设计:(1) 改进的均值流(iMF)动作生成器,将昂贵的多步扩散采样简化为1-3步动作生成;(2) 注意力残差(AttnRes),一种动态深度可分特征路由机制,替代均匀残差累积,更好保留任务相关多模态表征。我们在大规模仿真基准(LIBERO、RoboIMI)及真实机器人操作任务上评估ReactVLA。实验表明,它在性能上持续优于同类规模的VLA基线(如SmolVLA和π₀)。在高精度操作任务中,任务性能最高提升1.65倍,推理速度提升超4倍。最终,其真实策略延迟降至38.6毫秒以下,实现在物理机器人平台上的快速反应控制。

原文摘要 · Abstract (English)

Diffusion-based Vision-Language-Action (VLA) policies have demonstrated strong capability in modeling expressive and multimodal action distributions. However, their reliance on iterative sampling introduces substantial inference latency, which limits their applicability to reactive closed-loop robot manipulation. To address this limitation, we propose \texttt{ReactVLA}, a lightweight and low-latency VLA framework for real-time robotic manipulation. \texttt{ReactVLA} combines two complementary designs: (1) an improved Mean Flow (iMF) action generator that reduces expensive multi-step diffusion sampling to one-to-few-step action generation, and (2) Attention Residuals (AttnRes), a dynamic depth-wise feature routing mechanism that replaces uniform residual accumulation to better preserve task-relevant multimodal representations. We evaluate \texttt{ReactVLA} on large-scale simulation benchmarks, including LIBERO and RoboIMI, as well as real-world robotic manipulation tasks. Experimental results show that \texttt{ReactVLA} consistently outperforms similarly sized VLA baselines, including SmolVLA and $π_0$. On challenging precision manipulation tasks, \texttt{ReactVLA} achieves up to a 1.65$\times$ improvement in task performance while providing more than a 4$\times$ increase in inference speed compared with leading VLA models. Finally, it reduces real-world policy latency to below 38.6 ms, enabling fast reactive control on physical robot platforms. Please check out our project website at: https://game-loader.github.io/ReactVLA/.

机器人操作扩散模型实时控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。