arXiv:2510.15510cs.CVcs.RO2025-10中稿 · CVPR

用可学习提示让扩散模型适应机器人控制,性能超越现有方法

Exploring Conditions for Diffusion models in Robotic Control

  • 设计可学习任务提示和帧级视觉提示,实现动态适配
  • 在多个机器人控制基准上达到新最优,显著超越已有方法
  • 适合研究视觉表示与机器人控制融合的学者参考

尽管预训练视觉表征已显著推动模仿学习,但其通常为任务无关且在策略学习中保持冻结。本文探索利用预训练文本到图像扩散模型获取任务自适应视觉表征,而不微调模型本身。然而,我们发现直接应用文本条件——在其他视觉领域成功的方法——在控制任务中仅带来微弱甚至负面效果。这归因于扩散模型训练数据与机器人控制环境间的领域差异,因此我们主张采用考虑具体动态视觉信息的条件。为此,提出ORCA,引入可学习任务提示以适应控制环境,并使用捕捉精细帧级细节的视觉提示。通过新设计的条件促进任务自适应表征,本方法在多个机器人控制基准上取得当前最优性能,显著优于先前方法。

原文摘要 · Abstract (English)

While pre-trained visual representations have significantly advanced imitation learning, they are often task-agnostic as they remain frozen during policy learning. In this work, we explore leveraging pre-trained text-to-image diffusion models to obtain task-adaptive visual representations for robotic control, without fine-tuning the model itself. However, we find that naively applying textual conditions - a successful strategy in other vision domains - yields minimal or even negative gains in control tasks. We attribute this to the domain gap between the diffusion model's training data and robotic control environments, leading us to argue for conditions that consider the specific, dynamic visual information required for control. To this end, we propose ORCA, which introduces learnable task prompts that adapt to the control environment and visual prompts that capture fine-grained, frame-specific details. Through facilitating task-adaptive representations with our newly devised conditions, our approach achieves state-of-the-art performance on various robotic control benchmarks, significantly surpassing prior methods.

扩散模型机器人控制视觉表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。