用单图生成新视角,让机器人学会不依赖固定视角的抓取动作。
View-Invariant Policy Learning via Zero-Shot Novel View Synthesis
- 基于单张图像合成不同视角画面,实现零样本跨视角泛化。
- 在模拟和真实场景中,新视角下任务成功率提升显著超过基线。
- 适合需要多角度适应的机器人操控任务,尤其适用于少标注数据场景。
大规模视觉-运动策略学习是构建通用操作系统的有前途方法,但能在多种机器人形态、环境及观测模态下部署的策略仍难以实现。本文研究如何利用世界级大规模视觉数据,解决可泛化操作中的一个关键变量:观测视角差异。具体而言,我们考察单图像新视角合成模型,该模型通过给定单张输入图像,学习3D感知的场景先验,以渲染同一场景的其他视角图像。为实际应用于多样化机器人数据,这些模型需实现零样本运行,即在未见过的任务和环境中进行视角合成。我们通过名为视图合成增强(VISTA)的简单数据增强方案,实证分析了此类模型在从单视角演示数据中学习视角不变策略的能力。评估表明,采用本方法训练的策略在分布外相机视角下的鲁棒性显著优于基线,在模拟与真实世界操作任务中表现更优。视频与附加可视化见 https://s-tian.github.io/projects/vista。
原文摘要 · Abstract (English)
Large-scale visuomotor policy learning is a promising approach toward developing generalizable manipulation systems. Yet, policies that can be deployed on diverse embodiments, environments, and observational modalities remain elusive. In this work, we investigate how knowledge from large-scale visual data of the world may be used to address one axis of variation for generalizable manipulation: observational viewpoint. Specifically, we study single-image novel view synthesis models, which learn 3D-aware scene-level priors by rendering images of the same scene from alternate camera viewpoints given a single input image. For practical application to diverse robotic data, these models must operate zero-shot, performing view synthesis on unseen tasks and environments. We empirically analyze view synthesis models within a simple data-augmentation scheme that we call View Synthesis Augmentation (VISTA) to understand their capabilities for learning viewpoint-invariant policies from single-viewpoint demonstration data. Upon evaluating the robustness of policies trained with our method to out-of-distribution camera viewpoints, we find that they outperform baselines in both simulated and real-world manipulation tasks. Videos and additional visualizations are available at https://s-tian.github.io/projects/vista.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。