arXiv:2503.16776cs.CV2025-03被引 4

用视觉语言模型分析城市环境,实现零样本城市级智能评估

OpenCity3D: What do Vision-Language Models know about Urban Environments?

  • 基于航拍图像重建3D城市,用VLM完成高阶任务
  • 零样本/少样本下准确预测人口密度、房价、犯罪率等
  • 适合城市规划、政策制定与环境监测人员使用

视觉语言模型(VLMs)在3D场景理解中展现出巨大潜力,但主要应用于室内空间或自动驾驶,聚焦于分割等底层任务。本文将VLM扩展至城市尺度环境,利用多视角航拍图像生成的3D重建数据,提出OpenCity3D方法,解决人口密度估计、建筑年龄分类、房产价格预测、犯罪率评估和噪声污染评价等高层任务。实验表明,OpenCity3D具备出色的零样本与少样本能力,可适应新场景。本研究建立了语言驱动的城市分析新范式,适用于城市规划、政策制定与环境监测等领域。项目页面:opencity3d.github.io

原文摘要 · Abstract (English)

Vision-language models (VLMs) show great promise for 3D scene understanding but are mainly applied to indoor spaces or autonomous driving, focusing on low-level tasks like segmentation. This work expands their use to urban-scale environments by leveraging 3D reconstructions from multi-view aerial imagery. We propose OpenCity3D, an approach that addresses high-level tasks, such as population density estimation, building age classification, property price prediction, crime rate assessment, and noise pollution evaluation. Our findings highlight OpenCity3D's impressive zero-shot and few-shot capabilities, showcasing adaptability to new contexts. This research establishes a new paradigm for language-driven urban analytics, enabling applications in planning, policy, and environmental monitoring. See our project page: opencity3d.github.io

城市分析视觉语言模型零样本学习3D重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。