arXiv:2505.07819cs.ROcs.AI2025-05被引 13

通过三层层次化结构,让机器人视觉与动作生成更紧密协同。

H$^3$DP: Triply-Hierarchical Diffusion Policy for Visuomotor Learning

  • 分三层整合视觉与动作:深度感知输入、多尺度特征、分层扩散生成
  • 模拟任务平均性能提升27.5%,4个双臂真实任务表现优异
  • 适合做复杂视觉动作学习的科研人员和工程师

视觉运动策略学习在机器人操作中取得显著进展,现有方法多依赖生成模型建模动作分布,但常忽视视觉感知与动作预测间的紧密耦合。本文提出三重层次化扩散策略(H$^3$DP),显式引入三级层次结构以增强视觉特征与动作生成的融合。第一层为基于深度信息的RGB-D观测分层;第二层为多尺度视觉表征,编码不同粒度的语义特征;第三层为分层条件扩散过程,使粗粒度到细粒度动作的生成与对应视觉特征对齐。大量实验表明,H$^3$DP在44个仿真任务上相较基线平均提升27.5%;在4个挑战性双臂真实世界操作任务中也表现卓越。

原文摘要 · Abstract (English)

Visuomotor policy learning has witnessed substantial progress in robotic manipulation, with recent approaches predominantly relying on generative models to model the action distribution. However, these methods often overlook the critical coupling between visual perception and action prediction. In this work, we introduce $\textbf{Triply-Hierarchical Diffusion Policy}~(\textbf{H$^{\mathbf{3}}$DP})$, a novel visuomotor learning framework that explicitly incorporates hierarchical structures to strengthen the integration between visual features and action generation. H$^{3}$DP contains $\mathbf{3}$ levels of hierarchy: (1) depth-aware input layering that organizes RGB-D observations based on depth information; (2) multi-scale visual representations that encode semantic features at varying levels of granularity; and (3) a hierarchically conditioned diffusion process that aligns the generation of coarse-to-fine actions with corresponding visual features. Extensive experiments demonstrate that H$^{3}$DP yields a $\mathbf{+27.5\%}$ average relative improvement over baselines across $\mathbf{44}$ simulation tasks and achieves superior performance in $\mathbf{4}$ challenging bimanual real-world manipulation tasks. Project Page: https://lyy-iiis.github.io/h3dp/.

视觉动作扩散模型机器人学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。