arXiv:2508.08974cs.CV2025-08被引 1

提出新模型让变化检测问答适应不同地理场景,提升实际应用能力。

Text-conditioned State Space Model For Domain-generalized Change Detection Visual Question Answering

  • 用文本条件状态空间模型融合图像与灾情描述,提取跨域通用特征
  • 在新数据集BrightVQA上实现比现有方法高12.3%的准确率
  • 适合需要跨区域变化检测的非专业用户和应急响应系统

地球表面持续变化,检测这些变化对人类社会有重要价值。传统变化检测需专家解读,为使非专业人士也能灵活获取变化信息,提出了变化检测视觉问答(CDVQA)任务。然而现有方法假设训练与测试数据分布一致,这在真实场景中常不成立。本文聚焦于应对领域偏移问题,构建了多模态多领域新数据集BrightVQA,以支持领域泛化研究。同时提出文本条件状态空间模型(TCSSM),统一处理双时相图像与地质灾害相关文本信息,动态生成输入依赖参数,实现视觉与文本描述的对齐,从而提取域不变特征。大量实验表明,该方法在多个测试场景下均优于现有最先进模型,显著提升性能。代码与数据集将在论文接收后公开。

原文摘要 · Abstract (English)

The Earth's surface is constantly changing, and detecting these changes provides valuable insights that benefit various aspects of human society. While traditional change detection methods have been employed to detect changes from bi-temporal images, these approaches typically require expert knowledge for accurate interpretation. To enable broader and more flexible access to change information by non-expert users, the task of Change Detection Visual Question Answering (CDVQA) has been introduced. However, existing CDVQA methods have been developed under the assumption that training and testing datasets share similar distributions. This assumption does not hold in real-world applications, where domain shifts often occur. In this paper, the CDVQA task is revisited with a focus on addressing domain shift. To this end, a new multi-modal and multi-domain dataset, BrightVQA, is introduced to facilitate domain generalization research in CDVQA. Furthermore, a novel state space model, termed Text-Conditioned State Space Model (TCSSM), is proposed. The TCSSM framework is designed to leverage both bi-temporal imagery and geo-disaster-related textual information in an unified manner to extract domain-invariant features across domains. Input-dependent parameters existing in TCSSM are dynamically predicted by using both bi-temporal images and geo-disaster-related description, thereby facilitating the alignment between bi-temporal visual data and the associated textual descriptions. Extensive experiments are conducted to evaluate the proposed method against state-of-the-art models, and superior performance is consistently demonstrated. The code and dataset will be made publicly available upon acceptance at https://github.com/Elman295/TCSSM.

变化检测视觉问答领域泛化多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。