arXiv:2606.24208cs.RO2026-06

用物理约束优化扩散模型生成的机器人动作,提升可行性与成功率。

Grounding Generative Policies in Physics: Optimization-Guided Diffusion for Robot Control

论文配图:Grounding Generative Policies in Physics: Optimization-Guided Diffusion for Robot Control
图 1 · 摘自论文原文
  • 将扩散采样过程改为带约束的优化问题,动态修正生成结果。
  • 在灵巧抓取和动态操控任务中,成功率最高提升23个百分点。
  • 无需重训练模型,即可适配不同机器人,适合高要求控制场景。

扩散模型能从高维多模态分布中有效采样,但生成结果可能违反部署约束。对于任务空间的机器人策略,生成的抓取、路径点或轨迹虽分布合理,却可能不可达、碰撞或无法闭环执行,导致零样本跨机器人部署失败。本文提出一种推理时优化框架,通过将扩散引导建模为带约束的优化问题,使行为生成与物理可行性耦合。核心思想是用优化修正替代后向过程中的采样扰动,可在不重训练模型的前提下施加硬约束或软惩罚,同时保持样本接近学习到的先验分布。我们在灵巧抓取(需满足可达性与避障)和动态操控(需控制器可跟踪)任务上评估该方法。结果表明,优化引导去噪在各类机器人形态下,可行性媲美投影与梯度引导基线,同时更优地保持抓取质量,显著提升控制器可执行性与任务成功率,灵巧抓取任务成功率最高提升20个百分点,视觉-运动操控任务提升23个百分点。

原文摘要 · Abstract (English)

Diffusion models sample effectively from high-dimensional, multimodal distributions, but their outputs may violate deployment constraints. For task-space robot policies, generated grasps, waypoints, or trajectories can be distributionally valid yet infeasible, violating reachability, collision-avoidance, or closed-loop executability requirements. This embodiment gap limits zero-shot deployment across robots, even when the task-space behavior itself is transferable. We propose an inference-time optimization framework that couples the behavior generation to physical feasibility by formulating diffusion guidance as a constrained optimization problem. Our key insight is to replace the sampling perturbation in the backward process with an optimized correction, allowing hard constraints or soft penalties to be imposed during sampling without the need to retrain the diffusion model, while keeping samples close to the learned prior. We evaluate the method on dexterous grasp synthesis with reachability and collision-avoidance constraints, and dynamic manipulation with controller-level trackability constraints. Across settings and robot embodiments, optimization-guided denoising matches the feasibility of projection- and gradient-guidance baselines while better preserving grasp quality, and improving controller-level executability and task success, with task success improving by up to 20pp. on dexterous grasping and 23pp. on visuomotor manipulation over the best baseline.

机器人控制扩散模型物理约束优化引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。