arXiv:2608.05579cs.RO2026-08

用3D视觉模型统一机器人视角,让机械臂更懂空间位置。

ARGUS: Aligning Robot Scene Geometry Under Shifting Views with Large 3D Vision Models

论文配图:ARGUS: Aligning Robot Scene Geometry Under Shifting Views with Large 3D Vision Models
图 1 · 摘自论文原文
  • 用大模型将不同摄像头视角的图像对齐到统一坐标系。
  • 在多样视角数据上训练,成功率提升4-6倍,收敛更快。
  • 适合大规模多视角机器人数据集,提升泛化能力。

大规模视觉-运动策略在多种机器人操作任务中表现优异,但往往将场景几何与视角绑定,学习的是物体在图像中的位置而非任务空间中的真实位置,限制了从多视角数据(如 DROID、BridgeV2)中学习和跨视角泛化的能力。本文提出 ARGUS,一种观察预处理流程,利用大规模3D视觉模型将任意相机视角下的图像观测对齐至标准视角,再输入下游视觉-运动策略。在具有不同视角多样性的训练数据集上实验表明,该方法在有限视角和多视角训练场景下均优于先前方法。效率对比显示,ARGUS 能够更高效地利用多视角数据,在4-6倍更快的收敛速度下达到高成功率。结果表明,借助大规模3D视觉模型可减轻视觉-运动策略的学习负担,实现更高效的规模化多视角机器人数据学习。

原文摘要 · Abstract (English)

Large-scale visuomotor policies have demonstrated impressive performance across a wide range of robot manipulation tasks. However, despite this success, manipulation polices often entangle scene geometry with the corresponding viewpoint, learning where objects lie in an image rather than where it lies in the task space. This entanglement inherently limits the corresponding policy's ability to learn from viewpoint-diverse datasets (ex. DROID, BridgeV2) and generalize beyond the viewpoints captured in their training data. In this work, we present ARGUS, an observation pre-processing pipeline that uses large-scale 3D vision models to align image observations from arbitrary camera viewpoints into a canonical viewpoint before passing it to downstream visuomotor policies. Experiments across training datasets with varying levels of viewpoint diversity, from fixed multi-view camera configurations to highly varied camera placements, show that our method consistently outperforms prior approaches across both limited-view and view-diverse training regimes. In efficiency comparisons, ARGUS demonstrates an ability to learn from view-diverse data, converging to high success rates 4-6x faster than previous methods by leveraging a simplified observation space. Overall, our findings show that leveraging large-scale 3D vision models reduces the learning burden on visuomotor policies, enabling more efficient learning from large-scale, viewpoint-diverse robot datasets.

机器人3D视觉多视角对齐策略学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。