arXiv:2503.00339cs.RO2025-03ICML被引 11

通过复用历史动作减少采样步骤,实现视觉运动策略的快速推理。

Falcon: Fast Visuomotor Policies via Partial Denoising

  • 利用动作的时序依赖性,复用部分去噪的历史动作。
  • 在48个仿真环境和2个真实机器人实验中提速2-7倍。
  • 无需训练,可作为插件提升现有加速方法效率。

扩散策略在复杂视觉运动任务中广泛应用,因其能捕捉多模态动作分布。然而,生成动作所需的多次采样严重损害实时推理效率,限制其在实时决策场景中的应用。现有加速技术要么需要重新训练,要么在低采样步数下性能下降。本文提出Falcon,缓解速度与性能的权衡并实现进一步加速。核心思路是:视觉运动任务具有动作间的时序依赖性。Falcon通过复用历史信息中的部分去噪动作,而非每步从高斯噪声采样,结合当前观测,减少采样步数同时保持性能。重要的是,Falcon是无训练算法,可作为插件集成到现有加速技术上,提升决策效率。我们在48个仿真环境和2个真实机器人实验中验证,实现2-7倍加速,性能几乎无损,为高效视觉运动策略设计提供新方向。

原文摘要 · Abstract (English)

Diffusion policies are widely adopted in complex visuomotor tasks for their ability to capture multimodal action distributions. However, the multiple sampling steps required for action generation significantly harm real-time inference efficiency, which limits their applicability in real-time decision-making scenarios. Existing acceleration techniques either require retraining or degrade performance under low sampling steps. Here we propose Falcon, which mitigates this speed-performance trade-off and achieves further acceleration. The core insight is that visuomotor tasks exhibit sequential dependencies between actions. Falcon leverages this by reusing partially denoised actions from historical information rather than sampling from Gaussian noise at each step. By integrating current observations, Falcon reduces sampling steps while preserving performance. Importantly, Falcon is a training-free algorithm that can be applied as a plug-in to further improve decision efficiency on top of existing acceleration techniques. We validated Falcon in 48 simulated environments and 2 real-world robot experiments. demonstrating a 2-7x speedup with negligible performance degradation, offering a promising direction for efficient visuomotor policy design.

视觉运动扩散模型加速推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。