arXiv:2506.16201cs.ROcs.CV2025-06CVPR被引 17

用区域感知Mamba框架提升机器人操作的精度与速度

FlowRAM: Grounding Flow Matching Policy with Region-Aware Mamba Framework for Robotic Manipulation

  • 引入动态半径调度实现从全局到局部的自适应感知
  • 在RLBench上成功率达94.3%,比之前方法高12.0%
  • 4步内生成合理动作,推理速度显著提升

高精度机器人操作在众多工业和现实应用中至关重要,但现有基于扩散模型的策略学习方法因推理时迭代去噪过程导致计算效率低下,且未能充分挖掘生成模型在三维环境信息探索中的潜力。为此,我们提出FlowRAM框架,利用生成模型实现区域感知,支持高效多模态信息处理。具体地,设计动态半径调度,实现从全局场景理解到细粒度几何细节的自适应感知;引入状态空间模型融合多模态信息,同时保持线性计算复杂度;采用条件流匹配通过回归确定性向量场学习动作姿态,简化学习过程并维持性能。在标准操控基准RLBench上验证,FlowRAM达到最先进水平,尤其在高精度任务中平均成功率提升12.0%至94.3%。此外,可在少于4个时间步内生成物理合理的动作,显著提升推理速度。

原文摘要 · Abstract (English)

Robotic manipulation in high-precision tasks is essential for numerous industrial and real-world applications where accuracy and speed are required. Yet current diffusion-based policy learning methods generally suffer from low computational efficiency due to the iterative denoising process during inference. Moreover, these methods do not fully explore the potential of generative models for enhancing information exploration in 3D environments. In response, we propose FlowRAM, a novel framework that leverages generative models to achieve region-aware perception, enabling efficient multimodal information processing. Specifically, we devise a Dynamic Radius Schedule, which allows adaptive perception, facilitating transitions from global scene comprehension to fine-grained geometric details. Furthermore, we integrate state space models to integrate multimodal information, while preserving linear computational complexity. In addition, we employ conditional flow matching to learn action poses by regressing deterministic vector fields, simplifying the learning process while maintaining performance. We verify the effectiveness of the FlowRAM in the RLBench, an established manipulation benchmark, and achieve state-of-the-art performance. The results demonstrate that FlowRAM achieves a remarkable improvement, particularly in high-precision tasks, where it outperforms previous methods by 12.0% in average success rate. Additionally, FlowRAM is able to generate physically plausible actions for a variety of real-world tasks in less than 4 time steps, significantly increasing inference speed.

机器人操作生成模型状态空间模型高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。