arXiv:2502.09649cs.AIcs.CV2025-02被引 2

用视觉语言模型先验提升机器人在复杂干扰下的操作能力

ImitDiff: Transferring Foundation-Model Priors for Distraction Robust Visuomotor Policy

  • 通过双分辨率感知融合全局布局与局部细节,聚焦任务相关区域
  • 在复杂场景中优于现有方法,零样本泛化能力强
  • 动作生成速度提升十倍,适合实时机器人应用

视觉-运动模仿学习策略使机器人能高效从视觉示范中习得操作技能。然而,随着场景复杂度和视觉干扰增加,简单环境下表现良好的策略常出现显著性能下降。为此,我们提出ImitDiff,一种基于扩散模型、利用视觉-语言基础模型细粒度语义先验的模仿学习策略。该方法将高层指令转化为像素级视觉语义掩码,指导双分辨率感知流程:低分辨率捕捉整体布局等全局上下文,高分辨率提取几何细节等局部特征,使策略聚焦任务相关区域。此外,我们引入一致性驱动的扩散变压器动作头,实现视觉语义条件与实时动作生成之间的桥梁。大量实验表明,ImitDiff在复杂场景和视觉干扰下超越现有视觉-语言操控框架及视觉-运动模仿学习策略,尤其在涉及新物体和干扰的零样本设置中表现出强泛化能力。同时,其一致性驱动的动作头在保持竞争成功率的前提下,推理速度提升一个数量级。

原文摘要 · Abstract (English)

Visuomotor imitation learning policies enable robots to efficiently acquire manipulation skills from visual demonstrations. However, as scene complexity and visual distractions increase, policies that perform well in simple settings often experience substantial performance degradation. To address this challenge, we propose ImitDiff, a diffusion-based imitation learning policy guided by fine-grained semantics within a dual-resolution workflow. Leveraging pretrained priors of vision-language foundation models, our method transforms high-level instructions into pixel-level visual semantic masks. These masks guide a dual-resolution perception pipeline that captures both global context (e.g., overall layout) from low-resolution observation and fine-grained local features (e.g., geometric details) from high-resolution observation, enabling the policy to focus on task-relevant regions. Additionally, we introduce a consistency-driven diffusion transformer action head that bridges visual semantic conditions and real-time action generation. Extensive experiments demonstrate that ImitDiff outperforms state-of-the-art vision-language manipulation frameworks, as well as visuomotor imitation learning policies, particularly under increased scene complexity and visual distractions. Notably, ImitDiff exhibits strong generalization in zero-shot settings involving novel objects and visual distractions. Furthermore, our consistency-driven action head achieves an order-of-magnitude improvement in inference speed while maintaining competitive success rates.

机器人操作扩散模型视觉语言模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。