arXiv:2508.21565cs.CV2025-08ICCV被引 1

测试视觉语言模型对城市空间的理解能力,发现微调能显著提升复杂问题表现。

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images

  • 用合成数据+思维链监督微调模型,增强城市场景理解能力。
  • 微调后模型在否定、反事实等难题上性能大幅提升。
  • 适合关注城市感知、模型泛化与合成数据应用的研究者。

有效理解城市场景需要对物体、布局和深度线索进行细粒度的空间推理。然而,当前预训练于通用场景的视觉语言模型(VLMs)在城市领域的迁移能力尚未充分探索。为此,我们对比评估了三种现成的VLMs——BLIP-2、InstructBLIP和LLaVA-1.5——在零样本设置下的表现,并研究了使用针对城市场景的合成视觉问答(VQA)数据集进行微调的效果。该数据集基于街景图像的分割、深度与目标检测预测构建,每道问题均配以大语言模型生成的思维链(CoT)答案,实现分步推理监督。结果表明,尽管模型在零样本下表现尚可,但通过我们的合成CoT监督数据集微调后,性能显著提升,尤其在否定和反事实类问题上。本研究将城市空间推理引入VLM的新挑战,并展示了合成数据构建作为通用模型适配专业领域的一种可行路径。

原文摘要 · Abstract (English)

Effectively understanding urban scenes requires fine-grained spatial reasoning about objects, layouts, and depth cues. However, how well current vision-language models (VLMs), pretrained on general scenes, transfer these abilities to urban domain remains underexplored. To address this gap, we conduct a comparative study of three off-the-shelf VLMs-BLIP-2, InstructBLIP, and LLaVA-1.5-evaluating both zero-shot performance and the effects of fine-tuning with a synthetic VQA dataset specific to urban scenes. We construct such dataset from segmentation, depth, and object detection predictions of street-view images, pairing each question with LLM-generated Chain-of-Thought (CoT) answers for step-by-step reasoning supervision. Results show that while VLMs perform reasonably well in zero-shot settings, fine-tuning with our synthetic CoT-supervised dataset substantially boosts performance, especially for challenging question types such as negation and counterfactuals. This study introduces urban spatial reasoning as a new challenge for VLMs and demonstrates synthetic dataset construction as a practical path for adapting general-purpose models to specialized domains.

视觉语言模型城市理解合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。