构建动态城市理解基准,评估模型对长期城市变化的感知能力
DynamicVL: Benchmarking Multimodal Large Language Models for Dynamic City Understanding
- 设计多时相遥感图像基准框架,覆盖42个美国大城市
- 18个主流多模态模型在长期分析中表现不佳,尤其在定量任务
- 推出专用指令数据集,提升模型对城市动态的语言理解能力
多模态大语言模型在视觉理解方面表现卓越,但在长期地球观测分析中的应用仍有限,主要集中在单时相或双时相影像。为弥补这一空白,我们提出DVL-Suite,一个基于遥感影像的长期城市动态分析综合框架。该框架包含14,871张高分辨率(1.0米)多时相图像,覆盖美国42个主要城市,时间跨度为2005至2023年,分为DVL-Bench和DVL-Instruct两部分。DVL-Bench包含六项城市理解任务,涵盖像素级变化检测、区域级定量分析到场景级城市叙事,反映城市扩张/转型、灾害评估及环境挑战等多样动态。我们评估了18个先进MLLMs,发现其在长期时间理解与定量分析方面存在明显局限。为此,我们构建了专门用于指令微调的DVL-Instruct数据集,以增强模型在多时相地球观测中的能力。基于此,我们开发了DVLChat基线模型,支持图像级问答与像素级分割,通过语言交互实现对城市动态的全面理解。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in visual understanding, but their application to long-term Earth observation analysis remains limited, primarily focusing on single-temporal or bi-temporal imagery. To address this gap, we introduce DVL-Suite, a comprehensive framework for analyzing long-term urban dynamics through remote sensing imagery. Our suite comprises 14,871 high-resolution (1.0m) multi-temporal images spanning 42 major cities in the U.S. from 2005 to 2023, organized into two components: DVL-Bench and DVL-Instruct. The DVL-Bench includes six urban understanding tasks, from fundamental change detection (pixel-level) to quantitative analyses (regional-level) and comprehensive urban narratives (scene-level), capturing diverse urban dynamics including expansion/transformation patterns, disaster assessment, and environmental challenges. We evaluate 18 state-of-the-art MLLMs and reveal their limitations in long-term temporal understanding and quantitative analysis. These challenges motivate the creation of DVL-Instruct, a specialized instruction-tuning dataset designed to enhance models' capabilities in multi-temporal Earth observation. Building upon this dataset, we develop DVLChat, a baseline model capable of both image-level question-answering and pixel-level segmentation, facilitating a comprehensive understanding of city dynamics through language interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。