用单步反向神经网络加速机器人操作动作生成。
Invertible Neural Network Adapter for One-Step Flow Matching in Robot Manipulation

- 基于可逆神经网络的单步去噪生成动作,避免迭代计算。
- 实测推理延迟从110毫秒降至61毫秒,效率提升近50%。
- 适合需快速响应的视觉-语言-动作机器人系统部署。
本文提出一种用于通用机器人操作的可逆神经网络适配器,通过单步去噪过程,根据多模态观测(视觉、语言、本体感知)生成高维精确动作。基于流匹配框架,该适配器将动作生成轨迹约束在可逆潜在空间中,实现仅需一次推理即可高效生成高质量灵巧动作。相比传统迭代式流匹配策略,新方法显著降低推理复杂度,同时保持高精度与稳定性。在多种仿真基准和真实机器人平台上进行的大量实验表明:在仿真任务中表现优于或接近当前最优水平;在真实场景中,视觉-语言-动作模型平均推理延迟由110毫秒降至61毫秒,任务性能保持优异。
原文摘要 · Abstract (English)
This paper presents an invertible neural network adapter for general robotic manipulation, designed to generate precise high-dimensional actions conditioned on multimodal observations, including visual, linguistic, and proprioceptive inputs, through a one-step denoising process. Built upon a flow-matching formulation, the proposed adapter effectively constrains the action generation trajectory within an invertible latent space, thereby enabling efficient and high-quality dexterous action synthesis with only a single inference step. Compared with conventional iterative flow-matching policies, the proposed framework substantially reduces inference complexity while maintaining strong action prediction accuracy and stability. Extensive experiments are conducted across a diverse set of simulation benchmarks and real-world robotic platforms to evaluate the effectiveness of the proposed method. Across simulation benchmarks, the proposed adapter consistently demonstrates superior or near state-of-the-art performance on a wide range of manipulation tasks. Furthermore, real-world experiments reveal a significant improvement in inference efficiency for vision-language-action (VLA) models, reducing the average inference latency from 110 ms to 61 ms while maintaining strong task performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。