arXiv:2602.02396cs.ROcs.LG2026-02

PRISM实现单次生成的多模态机器人模仿学习,速度快且精度高。

PRISM: Performer RS-IMLE for Single-pass Multisensory Imitation Learning

  • 基于改进的IMLE方法,单次生成动作,避免迭代采样延迟。
  • 在真实机器人上成功率比扩散模型高10%-25%,控制频率达30-50Hz。
  • 适用于多传感器输入,适合复杂物理任务与大规模仿真测试。

机器人模仿学习需同时满足实时控制、多模态感知和动作分布建模需求。现有生成方法如扩散模型、流匹配和隐式最大似然估计(IMLE)常难以兼顾。本文提出PRISM,一种基于批量全局拒绝采样改进版IMLE的单次生成策略,结合时间多模态编码器(融合RGB、深度、触觉、音频与本体感知)和使用Performer架构的线性注意力生成器。在真实硬件平台验证,包括配备7-DoF机械臂的Unitree Go2和UR5机械臂,在预操作停车、高精度插入及多物体抓取等复杂任务中,成功率达扩散模型的10%-25%提升,同时保持30-50Hz闭环控制频率。在大型仿真基准(CALVIN、MetaWorld、Robomimic)上,于CALVIN(10%数据划分)中相较扩散模型提高约25%成功率,较流匹配提高约20%,轨迹抖动降低20至50倍。结果表明,PRISM兼具高速、高精度与多模态覆盖能力。

原文摘要 · Abstract (English)

Robotic imitation learning typically requires models that capture multimodal action distributions while operating at real-time control rates and accommodating multiple sensing modalities. Although recent generative approaches such as diffusion models, flow matching, and Implicit Maximum Likelihood Estimation (IMLE) have achieved promising results, they often satisfy only a subset of these requirements. To address this, we introduce PRISM, a single-pass policy based on a batch-global rejection-sampling variant of IMLE. PRISM couples a temporal multisensory encoder (integrating RGB, depth, tactile, audio, and proprioception) with a linear-attention generator using a Performer architecture. We demonstrate the efficacy of PRISM on a diverse real-world hardware suite, including loco-manipulation using a Unitree Go2 with a 7-DoF arm D1 and tabletop manipulation with a UR5 manipulator. Across challenging physical tasks such as pre-manipulation parking, high-precision insertion, and multi-object pick-and-place, PRISM outperforms state-of-the-art diffusion policies by 10-25% in success rate while maintaining high-frequency (30-50 Hz) closed-loop control. We further validate our approach on large-scale simulation benchmarks, including CALVIN, MetaWorld, and Robomimic. In CALVIN (10% data split), PRISM improves success rates by approximately 25% over diffusion and approximately 20% over flow matching, while simultaneously reducing trajectory jerk by 20x-50x. These results position PRISM as a fast, accurate, and multisensory imitation policy that retains multimodal action coverage without the latency of iterative sampling.

模仿学习多模态实时控制机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。