构建了最大规模的视觉语言空间推理数据集,助力模型理解复杂空间关系。
InternSpatial: A Comprehensive Dataset for Spatial Reasoning in Vision-Language Models
- 基于1200万问答对,覆盖单/多视角与19种指令格式。
- 在新基准上模型性能提升12.1%,多视角任务引入旋转角度预测新任务。
- 适合机器人、具身智能等需空间理解的场景研究者使用。
现有视觉语言模型的空间推理评测资源在规模、视觉多样性及指令表达力方面仍显不足。本文提出InternSpatial,目前最大的开源视觉语言空间推理数据集,包含1200万问答对,涵盖单视角与多视角设置,源自多样化视觉环境,并支持19种指令格式以反映不同查询风格。同时,我们构建了InternSpatial-Bench评估基准,用于单视角任务评估,并创新性地引入未被探索过的旋转角度预测任务,扩展多视角推理能力。实验表明,在InternSpatial上训练的模型在InternSpatial-Bench上提升12.1%,在VSI-Bench上提升10.7%,且在通用基准上保持优异表现。这些资源有望推动具身智能与机器人等实际应用中空间感知能力的发展。
原文摘要 · Abstract (English)
Recent benchmarks and datasets have been proposed to improve spatial reasoning in vision-language models (VLMs), yet existing open resources remain limited in scale, visual diversity, and instruction expressiveness. In this work, we introduce InternSpatial, the largest open-source dataset for spatial reasoning in VLMs, along with InternSpatial-Bench, a corresponding evaluation benchmark designed to assess spatial understanding under diverse instruction formats. InternSpatial comprises 12 million QA pairs spanning both single-view and multi-view settings, drawn from diverse visual environments and supporting 19 instruction formats that reflect varied query styles. For evaluation, we propose InternSpatial-Bench for single-view tasks and expand multi-view reasoning by introducing a novel rotation angle prediction task that has not been explored in prior work. Experimental results show that models trained on InternSpatial achieve 12.1% improvement on InternSpatial-Bench and 10.7% on VSI-Bench, while maintaining strong performance on general-purpose benchmarks. We hope these resources will support the development of spatially capable VLMs in practical applications such as robotics and embodied AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。