arXiv:2510.02268cs.ROcs.CV2025-10被引 15

让机器人模型知道摄像头位置,提升不同视角下的操作能力。

Do You Know Where Your Camera Is? View-Invariant Policy Learning with Camera Conditioning

  • 用光线的Plucker嵌入显式条件化相机外参,增强视角不变性。
  • 在多个任务上,带外参条件化的策略性能提升30%以上,且无需深度信息。
  • 适合做视觉控制、多视角模仿学习的研究者和开发者参考。

我们研究通过显式引入相机外参来实现视角不变的模仿学习。利用每像素光线的Plucker嵌入,发现对相机外参进行条件化能显著提升标准行为克隆策略(如ACT、Diffusion Policy和SmolVLA)在不同视角下的泛化能力。为评估策略在真实视角变化下的鲁棒性,我们在RoboSuite和ManiSkill中引入六个操作任务,包含“固定”与“随机”场景变体,将背景线索与相机位姿解耦。分析表明,未使用外参的策略常依赖固定场景中的静态背景视觉线索推断相机位姿;当工作区几何或相机位置改变时,该捷径失效。而加入外参条件化后,性能得以恢复,并实现仅用RGB图像的鲁棒控制,无需深度信息。相关任务、示范数据及代码已开源。

原文摘要 · Abstract (English)

We study view-invariant imitation learning by explicitly conditioning policies on camera extrinsics. Using Plucker embeddings of per-pixel rays, we show that conditioning on extrinsics significantly improves generalization across viewpoints for standard behavior cloning policies, including ACT, Diffusion Policy, and SmolVLA. To evaluate policy robustness under realistic viewpoint shifts, we introduce six manipulation tasks in RoboSuite and ManiSkill that pair "fixed" and "randomized" scene variants, decoupling background cues from camera pose. Our analysis reveals that policies without extrinsics often infer camera pose using visual cues from static backgrounds in fixed scenes; this shortcut collapses when workspace geometry or camera placement shifts. Conditioning on extrinsics restores performance and yields robust RGB-only control without depth. We release the tasks, demonstrations, and code at https://ripl.github.io/know_your_camera/ .

模仿学习视角不变视觉控制机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。