arXiv:2512.19347cs.RO2025-12中稿 · ICML被引 5

提出单步生成机器人操控新方法,兼顾精度与实时性。

OMP: One-step Meanflow Policy with Directional Alignment

  • 引入方向对齐机制,使预测速度与真实均值速度同步
  • 用微分导数方程近似雅可比-向量积,降低内存消耗
  • 在多个基准上实现高精度操控,且推理速度快

机器人操作越来越多采用数据驱动的生成策略,但存在持续权衡:扩散模型推理延迟高,而流模型常需复杂架构约束。尽管图像生成领域已有均值流(MeanFlow)实现单步推理,其直接应用于机器人时面临理论缺陷,如低速区的谱偏差和梯度饥饿。为此,我们提出单步均值流策略(OMP),专为高保真、实时操控设计。引入轻量级方向对齐机制,显式同步预测速度与真实均值速度;同时采用微分导数方程(DDE)近似雅可比-向量积(JVP),解耦前后向传播,显著降低内存复杂度。在Adroit和Meta-World基准上的大量实验表明,OMP在成功率和轨迹精度上优于现有最先进方法,尤其在高精度任务中表现突出,同时保持单步生成的高效性。

原文摘要 · Abstract (English)

Robot manipulation has increasingly adopted data-driven generative policy frameworks, yet the field faces a persistent trade-off: diffusion models suffer from high inference latency, while flow-based methods often require complex architectural constraints. Although in image generation domain, the MeanFlow paradigm offers a path to single-step inference, its direct application to robotics is impeded by critical theoretical pathologies, specifically spectral bias and gradient starvation in low-velocity regimes. To overcome these limitations, we propose the One-step MeanFlow Policy (OMP), a novel framework designed for high-fidelity, real-time manipulation. We introduce a lightweight directional alignment mechanism to explicitly synchronize predicted velocities with true mean velocities. Furthermore, we implement a Differential Derivation Equation (DDE) to approximate the Jacobian-Vector Product (JVP) operator, which decouples forward and backward passes to significantly reduce memory complexity. Extensive experiments on the Adroit and Meta-World benchmarks demonstrate that OMP outperforms state-of-the-art methods in success rate and trajectory accuracy, particularly in high-precision tasks, while retaining the efficiency of single-step generation.

机器人操控单步生成流模型实时控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。