arXiv:2412.09265cs.ROcs.LG2024-12被引 13

让扩散模型机器人策略提速6倍,同时保持高精度动作输出。

Score and Distribution Matching Policy: Advanced Accelerated Visuomotor Policies via Matched Distillation

  • 用两阶段优化将扩散模型转为单步生成器,提升推理速度。
  • 在57项仿真任务中实现6倍加速,动作质量达顶尖水平。
  • 双教师机制增强鲁棒性,适合高频实时控制场景。

基于扩散的视觉-运动策略虽能建模复杂机械臂轨迹,但推理时间长,难以满足高频控制需求。现有一致性蒸馏方法虽可提速,但引入误差影响动作质量。为此,我们提出评分与分布匹配策略(SDM Policy),通过两阶段优化将扩散策略转化为单步生成器:先进行评分匹配以对齐真实动作分布,再通过分布匹配最小化KL散度以保证一致性。采用双教师机制,冻结教师确保稳定性,非冻结教师用于对抗训练,提升鲁棒性与分布对齐能力。在包含57个任务的仿真基准上测试,该方法实现6倍推理加速,同时达到当前最优动作质量,为高频机器人任务提供高效可靠的解决方案。

原文摘要 · Abstract (English)

Visual-motor policy learning has advanced with architectures like diffusion-based policies, known for modeling complex robotic trajectories. However, their prolonged inference times hinder high-frequency control tasks requiring real-time feedback. While consistency distillation (CD) accelerates inference, it introduces errors that compromise action quality. To address these limitations, we propose the Score and Distribution Matching Policy (SDM Policy), which transforms diffusion-based policies into single-step generators through a two-stage optimization process: score matching ensures alignment with true action distributions, and distribution matching minimizes KL divergence for consistency. A dual-teacher mechanism integrates a frozen teacher for stability and an unfrozen teacher for adversarial training, enhancing robustness and alignment with target distributions. Evaluated on a 57-task simulation benchmark, SDM Policy achieves a 6x inference speedup while having state-of-the-art action quality, providing an efficient and reliable framework for high-frequency robotic tasks.

扩散模型机器人控制策略蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。