构建遥感变化检测新基准,支持描述变化和问答交互。
JL1-CC&QA: Extending the JL1-CD Benchmark with Change Captioning and Question Answering

- 基于吉林一号卫星5000对影像,新增变化描述与问答标注层。
- 含1.7万条高质量变化描述和2万条多类型问答对。
- 适合研究多任务遥感变化理解与人机交互的学者使用。
遥感变化检测传统上聚焦于像素级二值分割,仅定位变化位置而无法说明变化内容与原因。为弥合这一语义鸿沟,我们提出JL1-CC&QA,一个扩展自JL1-CD数据集的多任务基准,新增变化描述(CC)与变化问答(QA)两层标注。该基准基于吉林一号卫星获取的5,000对双时相影像(地面采样距离0.5–0.75米),包含:(i) JL1-CC,提供17,021条经质量验证的变化描述,涵盖多样化的地表覆盖转变;(ii) JL1-QA,包含20,060个跨八类问题的问答对,支持对地表变化的细粒度、交互式探查。所有标注通过三阶段流程生成:多模态大语言模型生成、视觉对齐的大语言模型判断、人工专家验证。我们期望该基准能作为统一融合二值变化掩码、变化描述与变化导向问答的公共资源,推动遥感领域多任务变化理解的发展。数据集可于https://github.com/circleLZY/JL1-CD 获取。
原文摘要 · Abstract (English)
Remote sensing change detection (CD) traditionally focuses on pixel-level binary segmentation, which identifies where changes occur but neither what nor why. To bridge this semantic gap, we introduce JL1-CC&QA, a multi-task benchmark that extends the JL1-CD dataset with two complementary annotation layers: change captioning (CC) and change question answering (QA). Built upon 5,000 bi-temporal image pairs acquired by the Jilin-1 satellite at 0.5-0.75m ground sample distance, the benchmark comprises: (i) JL1-CC, providing 17,021 quality-verified captions that describe diverse land-cover transformations; and (ii) JL1-QA, offering 20,060 question-answer pairs across eight question types, enabling fine-grained, interactive interrogation of surface changes. All annotations are produced via a three-stage pipeline consisting of multi-modal large language model (LLM) generation, vision-grounded LLM judging, and human expert verification. We hope that JL1-CC&QA, as a benchmark unifying binary change masks, change captions, and change-oriented QA over the same image set, will serve as a valuable resource for the community to advance multi-task change understanding in remote sensing. The dataset is available at https://github.com/circleLZY/JL1-CD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。