arXiv:2603.27967cs.CV2026-03中稿 · CVPR

构建跨视角关系数据集,提升视觉语言模型多视角空间推理能力

Learning Multi-View Spatial Reasoning from Cross-View Relations

  • 构建10万样本跨视角数据集,涵盖3类空间推理任务
  • 在多个基准上显著提升多视角推理性能,机器人任务成功率提高
  • 适合研究具身智能、机器人视觉与多视角理解的学者

视觉语言模型在单视角视觉任务中表现优异,但在理解三维环境和跨视角操作物体所需的多视角空间推理方面存在不足。本文提出跨视角关系(XVR)数据集,包含10万组来自1.8万个多样化3D场景和7万条机器人操作轨迹的视觉-问题-答案样本,覆盖三类基础空间推理任务:对应性(跨视图匹配物体)、验证性(判断空间关系)和定位性(确定物体位置)。在XVR上微调的视觉语言模型,在现有多个多视角与机器人空间推理基准(MindCube 和 RoboSpatial)上取得显著提升。当作为视觉语言动作模型的骨干网络时,XVR训练的表征使RoboCasa任务的成功率提高。结果表明,显式学习跨视角空间关系能有效增强多视角推理能力,并可迁移至真实机器人操作任务。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have achieved impressive results on single-view vision tasks, but lack the multi-view spatial reasoning capabilities essential for embodied AI systems to understand 3D environments and manipulate objects across different viewpoints. In this work, we introduce Cross-View Relations (XVR), a large-scale dataset designed to teach VLMs spatial reasoning across multiple views. XVR comprises 100K vision-question-answer samples derived from 18K diverse 3D scenes and 70K robotic manipulation trajectories, spanning three fundamental spatial reasoning tasks: Correspondence (matching objects across views), Verification (validating spatial relationships), and Localization (identifying object positions). VLMs fine-tuned on XVR achieve substantial improvements on established multi-view and robotic spatial reasoning benchmarks (MindCube and RoboSpatial). When integrated as backbones in Vision-Language-Action models, XVR-trained representations improve success rates on RoboCasa. Our results demonstrate that explicit training on cross-view spatial relations significantly enhances multi-view reasoning and transfers effectively to real-world robotic manipulation.

空间推理多视角理解机器人视觉视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。