提出细粒度遥感图像描述新任务,让模型像人一样精准描述地物特征与关系。
DescribeEarth: Describe Anything for Remote Sensing Images
- 设计对象级细粒度描述任务Geo-DLC,突破传统图像级描述局限。
- 构建包含26万实例的DE-Dataset,支持属性、关系与上下文多维度标注。
- 推出专用模型DescribeEarth,显著提升复杂场景下描述准确性与丰富性。
遥感图像的自动化文本描述对环境监测、城市规划和灾害管理等应用至关重要。然而,现有研究多停留在图像级别,缺乏对地物的细粒度解析,难以充分挖掘遥感图像中的语义与结构信息。为此,我们提出Geo-DLC任务——面向遥感图像的对象级细粒度描述。为此构建了大规模数据集DE-Dataset,涵盖25类、261,806个标注实例,提供对象属性、关系与上下文的详细描述。同时引入基于大语言模型的问答式评估框架DE-Benchmark,系统评测模型在Geo-DLC任务上的能力。我们还提出DescribeEarth,一种专为Geo-DLC设计的多模态大模型架构,融合尺度自适应焦点策略与领域引导融合模块,利用遥感视觉-语言模型特征,有效编码高分辨率细节与类别先验,同时保持全局上下文。DescribeEarth在DE-Benchmark上持续优于当前主流通用多模态大模型,尤其在捕捉内在地物特征与周边环境属性方面表现优异,适用于简单、复杂乃至分布外的遥感场景。所有数据、代码与权重已开源。
原文摘要 · Abstract (English)
Automated textual description of remote sensing images is crucial for unlocking their full potential in diverse applications, from environmental monitoring to urban planning and disaster management. However, existing studies in remote sensing image captioning primarily focus on the image level, lacking object-level fine-grained interpretation, which prevents the full utilization and transformation of the rich semantic and structural information contained in remote sensing images. To address this limitation, we propose Geo-DLC, a novel task of object-level fine-grained image captioning for remote sensing. To support this task, we construct DE-Dataset, a large-scale dataset contains 25 categories and 261,806 annotated instances with detailed descriptions of object attributes, relationships, and contexts. Furthermore, we introduce DE-Benchmark, a LLM-assisted question-answering based evaluation suite designed to systematically measure model capabilities on the Geo-DLC task. We also present DescribeEarth, a Multi-modal Large Language Model (MLLM) architecture explicitly designed for Geo-DLC, which integrates a scale-adaptive focal strategy and a domain-guided fusion module leveraging remote sensing vision-language model features to encode high-resolution details and remote sensing category priors while maintaining global context. Our DescribeEarth model consistently outperforms state-of-the-art general MLLMs on DE-Benchmark, demonstrating superior factual accuracy, descriptive richness, and grammatical soundness, particularly in capturing intrinsic object features and surrounding environmental attributes across simple, complex, and even out-of-distribution remote sensing scenarios. All data, code and weights are released at https://github.com/earth-insights/DescribeEarth.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。