arXiv:2606.14535cs.RO2026-06

用单个摄像头实现精准抓取,靠轨迹引导视觉注意力。

Spatially Conditioned Diffusion Policy: Learning Precise and Robust Manipulation with a Single RGB Camera

论文配图:Spatially Conditioned Diffusion Policy: Learning Precise and Robust Manipulation with a Single RGB Camera
图 1 · 摘自论文原文
  • 用机械臂轨迹作视觉注意力锚点,引导模型聚焦关键区域。
  • 仿真中超越主流单视角方法,接近多摄像头性能。
  • 适合无额外相机的机器人操作场景,抗干扰能力强。

近期视觉模仿学习系统普遍采用多摄像头配置,尤其是腕部安装的摄像头作为标准方案。然而,仅依靠单个全局视角进行操作仍具挑战性,因为策略需捕捉精细交互细节并识别任务相关区域,而无法依赖局部腕部视图。为此,我们提出空间条件扩散策略(SCDP),一种基于扩散模型的视觉-运动策略,在单摄像头设置下实现精确且鲁棒的操作。核心思路是:末端执行器轨迹可作为视觉注意力锚点,反映任务相关区域。SCDP包含两个关键组件:(i) 多尺度特征提取的视觉编码器,兼顾全局上下文与细粒度视觉特征;(ii) 空间条件模块,在扩散过程中沿中间末端执行器轨迹采样点级特征。大量仿真实验表明,SCDP持续优于强基准单视角方法,并达到与多摄像头基线相当的性能。真实世界实验进一步验证了其在视觉干扰下的精准操作与鲁棒性,凸显单摄像头模仿学习的潜力。

原文摘要 · Abstract (English)

Recent visual imitation learning systems have widely adopted multi-camera setups with wrist-mounted cameras as the de facto standard. However, manipulation from a single global view remains challenging, as the policy should capture fine-grained interaction details and identify task-relevant regions without local wrist views. To address this challenge, we present Spatially Conditioned Diffusion Policy (SCDP), a diffusion-based visuomotor policy that achieves precise and robust manipulation in a single-camera setting. Our key idea is that end-effector trajectories can serve as visual attention anchors that reflect task-relevant regions. Building on this idea, SCDP consists of two key components: (i) a visual encoder that produces multi-scale feature maps to capture both broader context and fine-grained visual features, and (ii) a spatial conditioning module that samples point-wise features along intermediate end-effector trajectories in the diffusion loop. Extensive simulation experiments show that SCDP consistently outperforms strong single-view baselines and achieves performance comparable to multi-camera baselines. Real-world experiments further demonstrate precise manipulation and robustness to visual distractors, highlighting the potential of single-camera imitation learning.

扩散模型机器人操作单摄像头

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。