arXiv:2606.00336cs.AIcs.LG2026-06

用可学习的连续参数控制扩散模型,实现行为精准调节。

From Noise to Control: Parameterized Diffusion Policies

论文配图:From Noise to Control: Parameterized Diffusion Policies
图 1 · 摘自论文原文
  • 在低维连续参数空间构建行为流形,让距离反映轨迹语义相似性
  • 相比标准扩散策略,新方法在复杂多模态任务中适应性显著提升
  • 适合需要生成新行为或快速调整约束的机器人控制场景

我们提出参数化扩散策略(PDP),一种基于可学习行为流形的扩散策略框架,该流形将低维连续参数嵌入其中。通过构建使潜在表示间距离反映物理轨迹语义相似性的流形,我们将扩散从随机多样性机制转变为可精确调控和优化的行为引导工具。该方法支持已知策略间的平滑插值,并可在不更新策略权重的情况下高效适应新约束。在模拟与真实机器人实验中,相比标准扩散策略,PDP在复杂多模态基准测试上显著提升了适应性能,尤其在需要合成新行为的场景中表现突出。

原文摘要 · Abstract (English)

We propose Parameterized Diffusion Policy (PDP), a framework for learning diffusion policies conditioned on low-dimensional, continuous parameters embedded in a learned behavior manifold. By constructing this manifold so that distances between latent representations reflect the semantic similarity between physical trajectories, we transform diffusion from a mechanism for stochastic diversity into a precise and optimizable tool for behavior steering. Our approach enables smooth interpolation between known strategies and efficient adaptation to novel constraints without updating policy weights. We demonstrate that PDP significantly improves adaptation performance on complex multimodal benchmarks in both simulated and real-robot experiments compared to standard diffusion policies, particularly in scenarios requiring the synthesis of novel behaviors.

扩散模型策略学习机器人控制行为生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。