arXiv:2503.12466cs.ROcs.CV2025-03中稿 · ICLR被引 5

用已有视觉模型组合生成更优策略,无需重新训练。

Modality-Composable Diffusion Policy via Inference-Time Distribution-level Composition

  • 通过组合多个预训练扩散策略的分布得分实现模态融合。
  • 在RoboTwin数据集上提升策略适应性与性能表现。
  • 适合需要跨模态、跨领域部署的机器人研究者使用。

扩散策略(Diffusion Policy, DP)因其能建模多分布动态而受到广泛关注,但现有方法通常依赖单一视觉模态(如RGB或点云),限制了准确性和泛化能力。尽管训练可处理异构多模态数据的通用DP能提升性能,却需巨大计算和数据开销。为此,我们提出一种新策略组合方法:利用基于不同视觉模态的多个预训练DP,通过其分布得分组合生成更具表达力的模态可组合扩散策略(MCDP),无需额外训练。在RoboTwin数据集上的大量实验证明,MCDP显著提升了策略的适应性与性能。该工作为现有DP的灵活组合提供了新思路,有助于推动跨模态、跨领域乃至跨机体的通用策略发展。代码已开源:https://github.com/AndyCao1125/MCDP。

原文摘要 · Abstract (English)

Diffusion Policy (DP) has attracted significant attention as an effective method for policy representation due to its capacity to model multi-distribution dynamics. However, current DPs are often based on a single visual modality (e.g., RGB or point cloud), limiting their accuracy and generalization potential. Although training a generalized DP capable of handling heterogeneous multimodal data would enhance performance, it entails substantial computational and data-related costs. To address these challenges, we propose a novel policy composition method: by leveraging multiple pre-trained DPs based on individual visual modalities, we can combine their distributional scores to form a more expressive Modality-Composable Diffusion Policy (MCDP), without the need for additional training. Through extensive empirical experiments on the RoboTwin dataset, we demonstrate the potential of MCDP to improve both adaptability and performance. This exploration aims to provide valuable insights into the flexible composition of existing DPs, facilitating the development of generalizable cross-modality, cross-domain, and even cross-embodiment policies. Our code is open-sourced at https://github.com/AndyCao1125/MCDP.

扩散模型机器人策略多模态策略组合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。