用合成数据提升视觉语言模型的空间推理能力
SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data
- 基于超详细图像描述生成340万条空间关系问答对
- 在What's Up基准上性能提升最高达49%
- 适合机器人、导航等需要空间理解的场景
视觉语言模型在图像描述、视觉问答等任务中表现良好,但在空间推理方面仍存在明显不足,这正是人类理解物理世界的关键能力。我们发现,现有主流视觉语言数据集中空间关系普遍稀少,仅有少数关系被充分覆盖,多数处于长尾分布。为填补这一空白,我们构建了一个聚焦空间推理的合成视觉问答数据集,数据来源于Localized Narratives、DOCCI和PixMo-Cap中的超详细图像描述,共包含45.5万样本、340万组问答对。以该数据集训练的SpaRE视觉语言模型,在空间推理基准测试中表现显著提升,尤其在What's Up基准上最高实现49%的性能增长,同时保持了通用任务的良好表现。本研究缩小了人类与视觉语言模型在空间推理能力上的差距,使模型在机器人、导航等现实应用中更具实用性。
原文摘要 · Abstract (English)
Vision-language models (VLMs) work well in tasks ranging from image captioning to visual question answering (VQA), yet they struggle with spatial reasoning, a key skill for understanding our physical world that humans excel at. We find that spatial relations are generally rare in widely used VL datasets, with only a few being well represented, while most form a long tail of underrepresented relations. This gap leaves VLMs ill-equipped to handle diverse spatial relationships. To bridge it, we construct a synthetic VQA dataset focused on spatial reasoning generated from hyper-detailed image descriptions in Localized Narratives, DOCCI, and PixMo-Cap. Our dataset consists of 455k samples containing 3.4 million QA pairs. Trained on this dataset, our Spatial-Reasoning Enhanced (SpaRE) VLMs show strong improvements on spatial reasoning benchmarks, achieving up to a 49% performance gain on the What's Up benchmark, while maintaining strong results on general tasks. Our work narrows the gap between human and VLM spatial reasoning and makes VLMs more capable in real-world tasks such as robotics and navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。