arXiv:2507.12795cs.CVcs.AI2025-07被引 2

构建首个多场景户外视觉语言模型,解决跨模态缺失问题。

City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete Learning

  • 用联合概率空间融合2D/3D多模态数据,支持模态缺失下的推理
  • 在420万图像与4811万点云上训练,问答任务平均提升18.14%
  • 适用于车、无人机、飞机、卫星等多视角户外场景理解

场景理解使智能体能够解析和认知环境。现有大型视觉-语言模型(LVLM)主要聚焦于室内家居任务,应用于大规模户外场景时存在两大局限:其一,户外场景涵盖多视角(如鸟瞰、地面视图)和多传感器的大规模环境,而现有模型多基于人眼视角的单一视觉模态;其二,缺乏多领域感知的户外数据,难以有效融合2D与3D视觉信息。为此,我们构建首个面向多域感知的户外场景理解数据集SVM-City,源自多尺度、多视角、多模态指令微调数据,包含420万张图像、4811万点云及56.7万组问答对,覆盖车辆、低空无人机、高空飞机与卫星。为应对模态缺失,提出不完整多模态学习框架,设计城市级视觉-语言模型City-VLM。通过构建联合概率分布空间实现多模态融合,而非直接拼接。在三个典型户外任务上的实验表明,City-VLM在问答任务上平均性能超越现有模型18.14%。方法展现出强泛化能力与实际应用潜力。

原文摘要 · Abstract (English)

Scene understanding enables intelligent agents to interpret and comprehend their environment. While existing large vision-language models (LVLMs) for scene understanding have primarily focused on indoor household tasks, they face two significant limitations when applied to outdoor large-scale scene understanding. First, outdoor scenarios typically encompass larger-scale environments observed through various sensors from multiple viewpoints (e.g., bird view and terrestrial view), while existing indoor LVLMs mainly analyze single visual modalities within building-scale contexts from humanoid viewpoints. Second, existing LVLMs suffer from missing multidomain perception outdoor data and struggle to effectively integrate 2D and 3D visual information. To address the aforementioned limitations, we build the first multidomain perception outdoor scene understanding dataset, named \textbf{\underline{SVM-City}}, deriving from multi\textbf{\underline{S}}cale scenarios with multi\textbf{\underline{V}}iew and multi\textbf{\underline{M}}odal instruction tuning data. It contains $420$k images and $4, 811$M point clouds with $567$k question-answering pairs from vehicles, low-altitude drones, high-altitude aerial planes, and satellite. To effectively fuse the multimodal data in the absence of one modality, we introduce incomplete multimodal learning to model outdoor scene understanding and design the LVLM named \textbf{\underline{City-VLM}}. Multimodal fusion is realized by constructing a joint probabilistic distribution space rather than implementing directly explicit fusion operations (e.g., concatenation). Experimental results on three typical outdoor scene understanding tasks show City-VLM achieves $18.14 \%$ performance surpassing existing LVLMs in question-answering tasks averagely. Our method demonstrates pragmatic and generalization performance across multiple outdoor scenes.

视觉语言模型多模态融合户外理解点云处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。