让遥感变化检测能回答用户问题并指出变化位置
Show Me What and Where has Changed? Question Answering and Grounding for Remote Sensing Change Detection
- 提出新任务CDQAG,结合问答与视觉定位
- 构建360K数据集,覆盖10类地物与8种问题类型
- 模型可同时输出文字答案和变化区域图示,适合需解释的场景
遥感变化检测旨在从不同时期的遥感数据中感知地表变化,并反馈给人类。然而,现有方法多仅关注变化区域检测,缺乏与用户交互以识别用户期望变化的能力。本文提出新任务——变化检测问答与定位(CDQAG),通过提供可解释的文字答案和直观的视觉证据,拓展传统变化检测。为此,我们构建首个CDQAG基准数据集QAG-360K,包含超过360,000个问答对与高质量视觉掩码,涵盖10类主要地表覆盖类型及8种综合问题类型,为遥感应用提供丰富多样数据。此外,我们提出VisTA,一种统一问答与定位任务的简单有效基线方法,同时输出视觉与文本答案。该方法在经典变化检测视觉问答(CDVQA)及所提CDQAG数据集上均取得当前最优性能。大量定性与定量实验为开发更优的CDQAG模型提供了重要启示,期望推动这一重要但研究不足领域的进一步发展。相关数据集与代码已开源。
原文摘要 · Abstract (English)
Remote sensing change detection aims to perceive changes occurring on the Earth's surface from remote sensing data in different periods, and feed these changes back to humans. However, most existing methods only focus on detecting change regions, lacking the capability to interact with users to identify changes that the users expect. In this paper, we introduce a new task named Change Detection Question Answering and Grounding (CDQAG), which extends the traditional change detection task by providing interpretable textual answers and intuitive visual evidence. To this end, we construct the first CDQAG benchmark dataset, termed QAG-360K, comprising over 360K triplets of questions, textual answers, and corresponding high-quality visual masks. It encompasses 10 essential land-cover categories and 8 comprehensive question types, which provides a valuable and diverse dataset for remote sensing applications. Furthermore, we present VisTA, a simple yet effective baseline method that unifies the tasks of question answering and grounding by delivering both visual and textual answers. Our method achieves state-of-the-art results on both the classic change detection-based visual question answering (CDVQA) and the proposed CDQAG datasets. Extensive qualitative and quantitative experimental results provide useful insights for developing better CDQAG models, and we hope that our work can inspire further research in this important yet underexplored research field. The proposed benchmark dataset and method are available at https://github.com/like413/VisTA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。