构建可解释的遥感视觉问答系统,提升模型决策透明度。
Checkmate: interpretable and explainable RSVQA is the endgame
- 设计新数据集Chessboard,含312万+问题,答案分布均衡。
- 模型能定位图像中与答案相关的具体区域,实现细粒度推理。
- 适合关注遥感智能分析可信性的研究人员使用。
遥感视觉问答(RSVQA)面临模型决策难以理解且缺乏视觉依据的问题。现有模型常因数据集偏差导致捷径学习,缺乏可解释性。本文提出新数据集Chessboard,包含3,123,253个问题和均衡的答案分布,每个答案对应图像中的一个或多个像素区域,支持细粒度视觉推理。基于此,我们开发了可解释模型Checkmate,能够识别影响其判断的关键图像区域。在多种模型架构上的实验证明,该方法显著提升了RVSQA系统的透明度,支持更可信的决策过程。
原文摘要 · Abstract (English)
Remote Sensing Visual Question Answering (RSVQA) presents unique challenges in ensuring that model decisions are both understandable and grounded in visual content. Current models often suffer from a lack of interpretability and explainability, as well as from biases in dataset distributions that lead to shortcut learning. In this work, we tackle these issues by introducing a novel RSVQA dataset, Chessboard, designed to minimize biases through 3'123'253 questions and a balanced answer distribution. Each answer is linked to one or more cells within the image, enabling fine-grained visual reasoning. Building on this dataset, we develop an explainable and interpretable model called Checkmate that identifies the image cells most relevant to its decisions. Through extensive experiments across multiple model architectures, we show that our approach improves transparency and supports more trustworthy decision-making in RSVQA systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。