arXiv:2505.12312cs.CVcs.AI2025-05被引 9

构建32万条视觉空间问答数据,提升机器人对三维环境的理解能力

Visuospatial Cognitive Assistant

  • 基于真实室内视频构建大规模视觉空间问答数据集
  • 在8个基准任务上超越现有模型,绝对距离误差降低26.1%
  • 首次提供可解释的空间推理链数据,适合研究具身智能与认知建模

基于视频的空间认知对机器人和具身人工智能至关重要,但当前视觉语言模型面临挑战。本文提出两个关键贡献:首先,构建了包含322,003个问答对的ViCA-322K数据集,源自ARKitScenes、ScanNet、ScanNet++等真实室内视频,支持基于3D元数据的查询与视频复杂推理;其次,基于该数据集微调得到ViCA-7B模型,在全部八个VSI-Bench任务中达到新最优性能,显著优于现有模型(如绝对距离指标提升26.1%)。为增强可解释性,我们构建了含显式推理链的ViCA-Thinking-2.68K数据集,并微调出可生成推理过程的ViCA-7B-Thinking模型。本工作强调专用数据的重要性,为改进时空建模提供新路径。所有资源已公开,以推动鲁棒视觉空间智能研究。

原文摘要 · Abstract (English)

Video-based spatial cognition is vital for robotics and embodied AI but challenges current Vision-Language Models (VLMs). This paper makes two key contributions. First, we introduce ViCA (Visuospatial Cognitive Assistant)-322K, a diverse dataset of 322,003 QA pairs from real-world indoor videos (ARKitScenes, ScanNet, ScanNet++), offering supervision for 3D metadata-grounded queries and video-based complex reasoning. Second, we develop ViCA-7B, fine-tuned on ViCA-322K, which achieves new state-of-the-art on all eight VSI-Bench tasks, outperforming existing models, including larger ones (e.g., +26.1 on Absolute Distance). For interpretability, we present ViCA-Thinking-2.68K, a dataset with explicit reasoning chains, and fine-tune ViCA-7B to create ViCA-7B-Thinking, a model that articulates its spatial reasoning. Our work highlights the importance of targeted data and suggests paths for improved temporal-spatial modeling. We release all resources to foster research in robust visuospatial intelligence.

视觉空间具身智能多模态推理链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。