提出遥感图像复制移动伪造检测与问答新任务,提升对篡改场景的理解能力。
Copy-Move Forgery Detection and Question Answering for Remote Sensing Image
- 设计区域判别引导的多模态感知框架,融合源图与篡改图差异信息
- 构建5个覆盖29地区14国的遥感伪造数据集,涵盖真实性和类别平衡性
- 适合遥感安全、图像取证和多模态问答研究者使用
为满足土地资源监测与国防安全的实际需求,本文提出了遥感图像复制移动伪造问答(RSCMQA)任务。与传统遥感视觉问答(RSVQA)不同,RSCMQA聚焦于复杂篡改场景的解读及对象间关系推理。我们构建了一套全面的全球RSCMQA数据集,涵盖14个国家的29个地区。具体包括五个数据集:基础数据集RS-CMQA、类别平衡数据集RS-CMQA-B、高真实性数据集Real-RSCM、扩展数据集RS-TQA以及扩展类别平衡数据集RS-TQA-B。这些数据集填补了领域空白,兼具全面性、平衡性与挑战性。同时,提出一种区域判别引导的多模态复制移动伪造感知框架(CMFPF),通过提示源域与篡改域之间的差异与关联,显著提升对篡改图像问答的准确性。大量实验表明,该方法在RSCMQA任务上优于通用VQA与RSVQA模型。相关数据集与代码已公开于https://github.com/shenyedepisa/RSCMQA。
原文摘要 · Abstract (English)
Driven by practical demands in land resource monitoring and national defense security, this paper introduces the Remote Sensing Copy-Move Question Answering (RSCMQA) task. Unlike traditional Remote Sensing Visual Question Answering (RSVQA), RSCMQA focuses on interpreting complex tampering scenarios and inferring relationships between objects. We present a suite of global RSCMQA datasets, comprising images from 29 different regions across 14 countries. Specifically, we propose five distinct datasets, including the basic dataset RS-CMQA, the category-balanced dataset RS-CMQA-B, the high-authenticity dataset Real-RSCM, the extended dataset RS-TQA, and the extended category-balanced dataset RS-TQA-B. These datasets fill a critical gap in the field while ensuring comprehensiveness, balance, and challenge. Furthermore, we introduce a region-discrimination-guided multimodal copy-move forgery perception framework (CMFPF), which enhances the accuracy of answering questions about tampered images by leveraging prompt about the differences and connections between the source and tampered domains. Extensive experiments demonstrate that our method provides a stronger benchmark for RSCMQA compared to general VQA and RSVQA models. Our datasets and code are publicly available at https://github.com/shenyedepisa/RSCMQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。