arXiv:2605.01581cs.RO2026-05被引 2

用频域分析优化3D扩散策略,仅两步即可高效生成机器人动作。

Hyper-DP3: Frequency-Aware Right-Sizing of 3D Diffusion Policies for Visuomotor Control

论文配图:Hyper-DP3: Frequency-Aware Right-Sizing of 3D Diffusion Policies for Visuomotor Control
图 1 · 摘自论文原文
  • 从频域视角看动作轨迹,发现低频成分主导能量分布。
  • 理论证明只需两步去噪即达性能上限,误差由低频维度与高频残余决定。
  • 模型参数少于1%,推理延迟极低,适合真实机器人部署。

基于扩散的视觉-运动策略在机器人操作中表现良好,但现有方法仍沿用图像生成式的解码器和多步采样。本文从频域角度重新审视该设计:机器人动作轨迹高度平滑,其能量主要集中在少数低频离散余弦变换模式中。在此结构下,我们证明最优去噪器的误差受低频子空间维数和残余高频能量限制,表明去噪误差在极少数反向步骤后即饱和。这暗示动作去噪所需模型远比图像生成简单。基于此洞察,我们提出超轻量级3D扩散策略Hyper-DP3(HDP3),采用轻量级Diffusion Mixer解码器,支持两步DDIM推理。合成实验验证了理论并支持两步去噪的充分性。在RoboTwin2.0、Adroit、MetaWorld及真实任务中,HDP3以不足前序3D扩散策略1%的参数量,实现最先进性能,且推理延迟显著降低。

原文摘要 · Abstract (English)

Diffusion-based visuomotor policies perform well in robotic manipulation, yet current methods still inherit image-generation-style decoders and multi-step sampling. We revisit this design from a frequency-domain perspective. Robot action trajectories are highly smooth, with most energy concentrated in a few low-frequency discrete cosine transform modes. Under this structure, we show that the error of the optimal denoiser is bounded by the low-frequency subspace dimension and residual high-frequency energy, implying that denoising error saturates after very few reverse steps. This also suggests that action denoising requires a much simpler denoising model than image generation. Motivated by this insight, we propose Hyper-DP3 (HDP3), a pocket-scale 3D diffusion policy with a lightweight Diffusion Mixer decoder that supports two-step DDIM inference. Our synthetic experiments validate the theory and support the sufficiency of two-step denoising. Futhermore, across RoboTwin2.0, Adroit, MetaWorld, and real-world tasks, HDP3 achieves state-of-the-art performance with fewer than 1% of the parameters of prior 3D diffusion-based policies and substantially lower inference latency.

3D扩散机器人控制轻量化频域分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。