通过构建空间关系知识图谱,提升视觉语言模型在街景中的空间推理能力。
DM-KG: A Novel Method for Boosting Spatial Cognition of Vision-Language Models in Street View Imagery

- 从单张街景图中提取实体间方向与距离关系,构建结构化知识图谱。
- 距离估计误差降低31.1%,方向判断误差降低65.8%,问答准确率保持高位。
- 适合需精准空间认知的地理视觉问答任务,可解释且易于部署。
随着视觉语言模型(VLMs)在地理空间问答与视觉场景理解中的广泛应用,提升其在街景图像中进行复杂逻辑推理的空间认知能力已成为关键研究方向。然而,现有VLMs在感知真实街景中物体位置、距离和方向时,常出现“空间语义幻觉”,且此类错误难以追溯与校准,严重制约其在地理空间任务中的实际应用。为此,本文提出DM-KG(Direction-Metric Knowledge Graph),一种基于结构化的街景空间表征框架。该框架通过融合全景分割与度量深度估计,鲁棒地计算实体级3D空间坐标,并将实体对间的钟向角与欧氏距离编码为JSON格式的知识图谱,作为显式几何先验注入VLM以引导空间推理。在公开空间问答(QA)基准上的实验表明,DM-KG使距离估计的平均绝对误差(MAE)降低31.1%,方向判断的平均角度误差降低65.8%,同时维持高问答成功率。本研究构建了完整的增强推理流程,显著提升了VLM在街景场景下的空间认知能力,为开放环境中的地理视觉问答(GeoVQA)提供了一种灵活、通用且可解释的解决方案。
原文摘要 · Abstract (English)
As vision-language models (VLMs) are increasingly deployed in geospatial question answering and visual scene understanding, improving their spatial cognition capability on street view imagery for complex logical reasoning has emerged as a key research priority. However, existing VLMs frequently suffer from "spatial semantic hallucinations" when perceiving object locations, distances, and directions in real-world street view scenes. Furthermore, such errors are often recalcitrant to tracing and calibration, posing a critical bottleneck for their practical deployment in geospatial tasks. To address this pressing challenge, this study proposes DM-KG (Direction-Metric Knowledge Graph), a structurally grounded spatial representation framework for street view imagery. By explicitly extracting directional and metric relationships between entities from a single 2D image, this framework enhances the spatial reasoning accuracy of VLMs through a structured knowledge graph. Specifically, we integrate panoptic segmentation with metric depth estimation to robustly compute entity-level 3D spatial coordinates. Subsequently, we encode the clock azimuths and Euclidean distances of entity pairs into a JSON-formatted knowledge graph, which is injected into the VLM as an explicit geometric prior to guide spatial reasoning. Experimental results on public spatial question-answering (QA) benchmarks demonstrate that DM-KG reduces the mean absolute error (MAE) in distance estimation by 31.1% and the mean angular error in direction judgment by 65.8%, while simultaneously maintaining a high QA success rate. By establishing a complete, augmented reasoning pipeline, this research significantly improves the spatial cognitive capabilities of VLMs in street view scenarios, thereby providing a flexible, generalized, and interpretable framework for geographic visual question answering (GeoVQA) in open environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。