用遥感图像+大模型分析城市环境,辅助智慧城市建设。
Built Environment Reasoning from Remote Sensing Imagery Using Large Vision--Language Models

- 将多尺度遥感图像输入多模态大模型,进行环境推理。
- 对比InternVL和Qwen模型,验证其在城市建议中的准确性。
- 适用于城市规划、风险评估等智慧城市决策场景。
本研究探讨大型语言模型(LLMs)在智慧城市建设中的应用。核心思想是利用遥感影像表征建成环境,包括设计建议、可建性评估、土地利用模式及风险识别。我们以多空间尺度的遥感影像作为多模态语言建模的输入,评估其对建成环境相关推理的影响。同时,对比InternVL和Qwen等前沿大模型在生成环境建议时的准确性和可靠性。结果表明,将遥感影像与大语言模型结合,具有辅助智慧城市建设与决策的潜力。
原文摘要 · Abstract (English)
This work investigates the use of large language models (LLMs) for tasks in smart cities. The core idea is to leverage remote sensing imagery to characterize the built environment, including design suggestions, constructability assessment, landuse patterns, and risk identification. We examine remote sensing imagery at multiple spatial scales as inputs for multimodal language modeling and evaluate their effects on built-environment-related reasoning. In addition, we compare state-of-the-art LLMs, including InternVL and Qwen, in terms of accuracy and reliability when generating built environment recommendations. The results demonstrate the potential of integrating remote sensing imagery with large language models to assist smart cities and decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。