arXiv:2603.09415cs.ROcs.AI2026-03

用单步推理实现高帧率多模态轨迹生成,解决机器人操控延迟与分布坍塌问题。

From Flow to One Step: Real-Time Multi-Modal Trajectory Policies via Implicit Maximum Likelihood Estimation-based Distribution Distillation

  • 通过隐式最大似然估计蒸馏流模型,实现单次前向传播生成
  • 采用双向切比雪夫距离保持多模态分布,避免轨迹平均化
  • 融合多视角视觉与本体感知,支持实时重规划与动态扰动鲁棒性

基于扩散和流匹配的生成策略在机器人操作中表现优异,能建模多模态人类示范。但其依赖迭代常微分方程求解导致显著延迟,限制高频闭环控制。现有单步加速方法虽缓解开销,却常出现分布坍塌,生成平均化轨迹,无法执行连贯操作策略。本文提出一种框架,通过隐式最大似然估计(IMLE)将条件流匹配(CFM)教师模型蒸馏为快速单步学生模型。双向切比雪夫距离提供集合级目标,兼顾模式覆盖与保真度,使学生模型在单次前向传播中保留教师的多模态动作分布。统一感知编码器融合多视角RGB、深度图、点云及本体感知,生成几何感知表征。所提方法支持高频率控制,实现动态扰动下的实时滚动规划与更强鲁棒性。

原文摘要 · Abstract (English)

Generative policies based on diffusion and flow matching achieve strong performance in robotic manipulation by modeling multi-modal human demonstrations. However, their reliance on iterative Ordinary Differential Equation (ODE) integration introduces substantial latency, limiting high-frequency closed-loop control. Recent single-step acceleration methods alleviate this overhead but often exhibit distributional collapse, producing averaged trajectories that fail to execute coherent manipulation strategies. We propose a framework that distills a Conditional Flow Matching (CFM) expert into a fast single-step student via Implicit Maximum Likelihood Estimation (IMLE). A bi-directional Chamfer distance provides a set-level objective that promotes both mode coverage and fidelity, enabling preservation of the teacher multi-modal action distribution in a single forward pass. A unified perception encoder further integrates multi-view RGB, depth, point clouds, and proprioception into a geometry-aware representation. The resulting high-frequency control supports real-time receding-horizon re-planning and improved robustness under dynamic disturbances.

机器人控制多模态轨迹单步生成流匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。