arXiv:2510.26670cs.RO2025-10被引 3

让机器人快速学会多种操作,速度与多样性兼得

Hybrid Consistency Policy: Decoupling Multi-Modal Diversity and Real-Time Efficiency in Robotic Manipulation

  • 先短时随机去噪,再一步一致性跳跃生成动作
  • 25步+1跳达80步扩散模型精度,延迟大幅降低
  • 适合对实时性要求高的机器人控制场景

在视觉-运动策略学习中,基于扩散的模仿学习因其能捕捉多样化行为而广受欢迎。然而,传统基于普通和随机去噪的方法难以同时实现快速采样和强多模态。为此,我们提出混合一致性策略(HCP):先运行至自适应切换时间的短时随机前缀,再通过单步一致性跳跃生成最终动作。为对齐这一跳跃生成,HCP采用随时间变化的一致性蒸馏,结合轨迹一致性目标以保持相邻预测连贯性,以及去噪匹配目标以提升局部保真度。在仿真和真实机器人上,仅需25个SDE步骤加一次跳跃的HCP,即可达到80步DDPM教师模型的准确率和模式覆盖度,同时显著降低延迟。结果表明,多模态无需慢速推理,切换时间实现了模式保留与速度的解耦,为机器人策略提供了实用的精度-效率权衡。

原文摘要 · Abstract (English)

In visuomotor policy learning, diffusion-based imitation learning has become widely adopted for its ability to capture diverse behaviors. However, approaches built on ordinary and stochastic denoising processes struggle to jointly achieve fast sampling and strong multi-modality. To address these challenges, we propose the Hybrid Consistency Policy (HCP). HCP runs a short stochastic prefix up to an adaptive switch time, and then applies a one-step consistency jump to produce the final action. To align this one-jump generation, HCP performs time-varying consistency distillation that combines a trajectory-consistency objective to keep neighboring predictions coherent and a denoising-matching objective to improve local fidelity. In both simulation and on a real robot, HCP with 25 SDE steps plus one jump approaches the 80-step DDPM teacher in accuracy and mode coverage while significantly reducing latency. These results show that multi-modality does not require slow inference, and a switch time decouples mode retention from speed. It yields a practical accuracy efficiency trade-off for robot policies.

机器人控制扩散模型实时生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。