arXiv:2601.10477cs.CVcs.AI2026-01被引 3

用视觉语言推理实现城市社会语义分割,突破传统模型局限

Urban Socio-Semantic Segmentation with Vision-Language Reasoning

  • 通过跨模态识别与多阶段推理模拟人类标注过程
  • 在SocioSeg数据集上显著优于现有模型,零样本泛化能力强
  • 适合城市规划、智能交通等需要理解社会空间的场景

城市是人类活动的中心,包含丰富的语义实体。从卫星图像中分割这些实体对下游应用至关重要。当前先进分割模型能可靠识别由物理属性定义的实体(如建筑、水体),但难以处理由社会属性定义的类别(如学校、公园)。本文通过视觉-语言模型推理实现社会语义分割。为此,我们构建了名为SocioSeg的新数据集,包含卫星影像、数字地图及层级结构化的社会语义实体像素级标签。同时提出新颖的视觉-语言推理框架SocioReasoner,模拟人类识别与标注社会语义实体的过程,结合跨模态识别与多阶段推理,并采用强化学习优化不可微过程,激发视觉-语言模型的推理能力。实验表明,该方法在多个指标上超越现有最先进模型,且具备强零样本泛化能力。数据集与代码已开源,许可协议为Apache License 2.0。

原文摘要 · Abstract (English)

As hubs of human activity, urban surfaces consist of a wealth of semantic entities. Segmenting these various entities from satellite imagery is crucial for a range of downstream applications. Current advanced segmentation models can reliably segment entities defined by physical attributes (e.g., buildings, water bodies) but still struggle with socially defined categories (e.g., schools, parks). In this work, we achieve socio-semantic segmentation by vision-language model reasoning. To facilitate this, we introduce the Urban Socio-Semantic Segmentation dataset named SocioSeg, a new resource comprising satellite imagery, digital maps, and pixel-level labels of social semantic entities organized in a hierarchical structure. Additionally, we propose a novel vision-language reasoning framework called SocioReasoner that simulates the human process of identifying and annotating social semantic entities via cross-modal recognition and multi-stage reasoning. We employ reinforcement learning to optimize this non-differentiable process and elicit the reasoning capabilities of the vision-language model. Experiments demonstrate our approach's gains over state-of-the-art models and strong zero-shot generalization. The dataset and code are open-sourced under the Apache License 2.0 at https://github.com/AMAP-ML/SocioReasoner.

社会语义分割视觉语言模型城市感知多模态推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。