arXiv:2411.14476cs.CLcs.AI2024-11被引 10

用多模态大模型分析街景图,精准预测城市环境指标。

StreetviewLLM: Extracting Geographic Information Using a Chain-of-Thought Multimodal Large Language Model

  • 结合街景、坐标和文本,用思维链推理提取地理信息。
  • 在7个全球城市中,对人口密度等5项指标预测更准。
  • 适合城市规划与环境监测人员参考使用。

地理空间预测在灾害管理、城市规划和公共健康等领域至关重要。传统机器学习方法在处理街景图像等非结构化多模态数据时存在局限。为此,我们提出StreetViewLLM,一种融合大语言模型与思维链推理的新型框架,整合街景图像、地理坐标及文本数据,提升地理空间预测的精度与粒度。通过检索增强生成技术,该方法增强地理信息提取能力,实现对城市环境的细粒度分析。模型已在香港、东京、新加坡、洛杉矶、纽约、伦敦和巴黎七座全球城市验证,显著优于基线模型,在人口密度、医疗可及性、归一化差异植被指数、建筑高度和不透水面等城市指标预测上表现优异。结果表明,StreetViewLLM在预测准确性与城市环境洞察力方面均有提升,为大语言模型在城市分析、规划决策、基础设施管理和环境监测中的应用开辟新路径。

原文摘要 · Abstract (English)

Geospatial predictions are crucial for diverse fields such as disaster management, urban planning, and public health. Traditional machine learning methods often face limitations when handling unstructured or multi-modal data like street view imagery. To address these challenges, we propose StreetViewLLM, a novel framework that integrates a large language model with the chain-of-thought reasoning and multimodal data sources. By combining street view imagery with geographic coordinates and textual data, StreetViewLLM improves the precision and granularity of geospatial predictions. Using retrieval-augmented generation techniques, our approach enhances geographic information extraction, enabling a detailed analysis of urban environments. The model has been applied to seven global cities, including Hong Kong, Tokyo, Singapore, Los Angeles, New York, London, and Paris, demonstrating superior performance in predicting urban indicators, including population density, accessibility to healthcare, normalized difference vegetation index, building height, and impervious surface. The results show that StreetViewLLM consistently outperforms baseline models, offering improved predictive accuracy and deeper insights into the built environment. This research opens new opportunities for integrating the large language model into urban analytics, decision-making in urban planning, infrastructure management, and environmental monitoring.

城市分析多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。