arXiv:2603.23447cs.CVcs.AI2026-03

让大模型理解城市级3D场景,支持复杂空间推理。

3DCity-LLM: Empowering Multi-modality Large Language Models for 3D City-scale Perception and Understanding

  • 分三路编码目标、关系与全局场景,实现精细感知
  • 构建120万样本高质量数据集,覆盖7类城市任务
  • 适合城市规划、自动驾驶等需空间理解的领域

尽管多模态大语言模型在物体中心或室内场景表现优异,但将其扩展至3D城市级环境仍面临巨大挑战。为此,我们提出3DCity-LLM,一个统一框架,用于3D城市级视觉-语言感知与理解。该框架采用粗到精的特征编码策略,包含三个并行分支:目标物体、物体间关系和全局场景。为支持大规模训练,我们构建了3DCity-LLM-1.2M数据集,包含约120万条高质量样本,覆盖七类代表性任务,从细粒度物体分析到多维度场景规划。该数据集严格控制质量,融合显式3D数值信息与多样化的用户导向模拟,提升城市场景问答的多样性与真实性。此外,我们采用基于文本相似性度量与大模型语义评估的多维评测协议,确保对所有方法的评估准确且全面。在两个基准上的大量实验表明,3DCity-LLM显著优于现有最先进方法,为推动空间推理与城市智能提供了有前景的方向。源代码与数据集见https://github.com/SYSU-3DSTAILab/3D-City-LLM。

原文摘要 · Abstract (English)

While multi-modality large language models excel in object-centric or indoor scenarios, scaling them to 3D city-scale environments remains a formidable challenge. To bridge this gap, we propose 3DCity-LLM, a unified framework designed for 3D city-scale vision-language perception and understanding. 3DCity-LLM employs a coarse-to-fine feature encoding strategy comprising three parallel branches for target object, inter-object relationship, and global scene. To facilitate large-scale training, we introduce 3DCity-LLM-1.2M dataset that comprises approximately 1.2 million high-quality samples across seven representative task categories, ranging from fine-grained object analysis to multi-faceted scene planning. This strictly quality-controlled dataset integrates explicit 3D numerical information and diverse user-oriented simulations, enriching the question-answering diversity and realism of urban scenarios. Furthermore, we apply a multi-dimensional protocol based on text-similarity metrics and LLM-based semantic assessment to ensure faithful and comprehensive evaluations for all methods. Extensive experiments on two benchmarks demonstrate that 3DCity-LLM significantly outperforms existing state-of-the-art methods, offering a promising and meaningful direction for advancing spatial reasoning and urban intelligence. The source code and dataset are available at https://github.com/SYSU-3DSTAILab/3D-City-LLM.

3D感知城市智能多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。