arXiv:2508.06266cs.RO2025-08被引 4

让机器人抓取更准更快:通过几何约束和智能初始动作提升扩散策略性能

ADPro: a Test-time Adaptive Diffusion Policy via Manifold-constrained Denoising and Task-aware Initialization for Robotic Manipulation

  • 引入几何流形约束,引导去噪沿任务相关路径进行
  • 用粗略配准生成初始动作,提速收敛并减少无效探索
  • 无需重训练,在真实场景中提升成功率与执行效率

扩散策略已成为机器人操作中强大的视觉-运动控制器,具备稳定训练和多模态动作建模能力。然而,现有方法通常将动作生成视为无约束去噪过程,忽略了关于几何结构和控制规律的先验知识。本文提出自适应扩散策略(ADP),一种测试时适配方法,引入两项关键归纳偏置:首先,嵌入几何流形约束,使去噪更新对齐于任务相关的子空间,利用末端执行器与目标场景间的相对位姿作为自然梯度方向,沿操作流形的测地线路径引导去噪;其次,为减少不必要的探索并加速收敛,提出解析引导的初始化:不从无信息先验采样,而是通过计算夹爪与目标场景间的粗略配准,生成结构化的初始噪声动作。ADP兼容预训练扩散策略且无需重训练,可在测试时适配具体任务,从而增强跨新任务和环境的泛化能力。在RLBench、CALVIN及真实数据集上的实验表明,ADPro(ADP的实现)提升了成功率、泛化性与采样效率,执行速度最快提升25%,成功率超过强基线9个百分点。

原文摘要 · Abstract (English)

Diffusion policies have recently emerged as a powerful class of visuomotor controllers for robot manipulation, offering stable training and expressive multi-modal action modeling. However, existing approaches typically treat action generation as an unconstrained denoising process, ignoring valuable a priori knowledge about geometry and control structure. In this work, we propose the Adaptive Diffusion Policy (ADP), a test-time adaptation method that introduces two key inductive biases into the diffusion. First, we embed a geometric manifold constraint that aligns denoising updates with task-relevant subspaces, leveraging the fact that the relative pose between the end-effector and target scene provides a natural gradient direction, and guiding denoising along the geodesic path of the manipulation manifold. Then, to reduce unnecessary exploration and accelerate convergence, we propose an analytically guided initialization: rather than sampling from an uninformative prior, we compute a rough registration between the gripper and target scenes to propose a structured initial noisy action. ADP is compatible with pre-trained diffusion policies and requires no retraining, enabling test-time adaptation that tailors the policy to specific tasks, thereby enhancing generalization across novel tasks and environments. Experiments on RLBench, CALVIN, and real-world datasets show that ADPro, an implementation of ADP, improves success rates, generalization, and sampling efficiency, achieving up to 25% faster execution and 9% points over strong diffusion baselines.

扩散模型机器人操作测试时适应动作生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。