用零样本SAM实现跨域鲁棒的场景变化检测
Towards Generalizable Scene Change Detection
- 零样本调用SAM,通过伪掩码生成与几何语义匹配实现无引导检测
- 在未知环境和不同时间条件下性能从77.6%降至4.6%,新方法提升30.0%
- 提出新基准ChangeVPR和评估协议,适合关注泛化能力的研究者
当前先进场景变化检测(SCD)方法在训练数据上表现优异,但在未见环境和不同时间条件下可靠性急剧下降:域内性能从77.6%降至8.0%,时间条件变化下更降至4.6%,凸显对泛化性SCD的需求。本文提出通用场景变化检测框架GeSCF,解决未见域性能与时间一致性问题。方法零样本利用预训练分割任意模型(SAM),设计初始伪掩码生成与几何-语义掩码匹配机制,将用户提示和单图分割无缝转化为一对输入的场景变化检测。同时构建通用场景变化检测(GeSCD)基准,包含新指标与评估协议,并引入挑战性图像对数据集ChangeVPR,涵盖城市、郊区、农村等多种环境。跨数据集实验证明,GeSCF在现有SCD数据集平均提升19.2%,在ChangeVPR上提升30.0%,接近翻倍领先性能。
原文摘要 · Abstract (English)
While current state-of-the-art Scene Change Detection (SCD) approaches achieve impressive results in well-trained research data, they become unreliable under unseen environments and different temporal conditions; in-domain performance drops from 77.6% to 8.0% in a previously unseen environment and to 4.6% under a different temporal condition -- calling for generalizable SCD and benchmark. In this work, we propose the Generalizable Scene Change Detection Framework (GeSCF), which addresses unseen domain performance and temporal consistency -- to meet the growing demand for anything SCD. Our method leverages the pre-trained Segment Anything Model (SAM) in a zero-shot manner. For this, we design Initial Pseudo-mask Generation and Geometric-Semantic Mask Matching -- seamlessly turning user-guided prompt and single-image based segmentation into scene change detection for a pair of inputs without guidance. Furthermore, we define the Generalizable Scene Change Detection (GeSCD) benchmark along with novel metrics and an evaluation protocol to facilitate SCD research in generalizability. In the process, we introduce the ChangeVPR dataset, a collection of challenging image pairs with diverse environmental scenarios -- including urban, suburban, and rural settings. Extensive experiments across various datasets demonstrate that GeSCF achieves an average performance gain of 19.2% on existing SCD datasets and 30.0% on the ChangeVPR dataset, nearly doubling the prior art performance. We believe our work can lay a solid foundation for robust and generalizable SCD research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。