arXiv:2608.01066cs.RO2026-08

让机器人在视角变化下更稳定操作,提升真实场景适应力

OC-VLA++: Monocular Geometry-Guided Cross-View Consistency for Viewpoint-Robust Robotic Manipulation

论文配图:OC-VLA++: Monocular Geometry-Guided Cross-View Consistency for Viewpoint-Robust Robotic Manipulation
图 1 · 摘自论文原文
  • 用几何关联视角对监督动作预测一致性
  • 未见视角表现显著提升,相机位移越大越稳定
  • 适合需要多视角泛化的机器人抓取任务

我们提出 OC-VLA++,用于在摄像头覆盖有限时实现视角鲁棒的机器人操作。虽然 OC-VLA 将动作锚定在相机坐标系以对齐动作监督与视觉观测,但仅靠相机空间锚定仍可能过拟合于训练中观察到的少数视角。OC-VLA++ 通过引入基于几何引导的成对视角监督和显式的跨视角动作等变性目标来解决该问题。给定同一操作场景从几何相关视角拍摄的成对观测,模型被训练为使它们在相机坐标系中的预测对应相同的机器人坐标系动作。该目标显式监督了动作预测在不同视角间的变换方式,而非依赖图像级增强。实验表明,在摄像头覆盖受限条件下,未见视角的泛化能力显著提升,且随着摄像头位移增加,性能下降更平缓。结果确立了跨视角动作等变性作为观测中心动作锚定的有效补充,有助于实际部署的鲁棒性。

原文摘要 · Abstract (English)

We propose OC-VLA++, an extension of OC-VLA for viewpoint generalization under limited camera coverage. While OC-VLA grounds robot actions in the camera coordinate system to align action supervision with visual observations, camera-space grounding alone can still overfit to the few viewpoints observed during training. OC-VLA++ addresses this limitation by introducing geometry-guided paired-view supervision and an explicit cross-view action-equivariance objective. Given paired observations of the same manipulation scene from geometrically related viewpoints, the model is trained such that their camera-space predictions correspond to the same robot-frame action. This objective explicitly supervises how action predictions should transform across viewpoints, rather than relying solely on image-level augmentation. Experiments demonstrate substantial improvements in unseen-view generalization under limited camera coverage, with performance degrading more gracefully under increasing camera displacement. These results establish cross-view action equivariance as an effective complement to observation-centric action grounding for robust real-world deployment.

机器人操作视角鲁棒动作等变

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。