arXiv:2509.02273cs.CV2025-09被引 2

用视觉语言模型提升遥感图像异常检测,少样本下表现更优

RS-OOD: A Vision-Language Augmented Framework for Out-of-Distribution Detection in Remote Sensing

  • 结合遥感场景与语义的双提示对齐机制,增强空间-语义一致性
  • 在多个遥感数据集上超越现有方法,少样本下仍保持高精度
  • 自训练机制自动挖掘伪标签,减少人工标注依赖,适合资源有限场景

遥感图像中的分布外(OOD)检测对自主监测、灾害响应和环境评估至关重要。尽管自然图像的OOD检测已取得显著进展,但现有方法和基准因数据稀缺、多尺度场景复杂及分布偏移严重,难以适用于遥感图像。为此,我们提出RS-OOD框架,利用面向遥感的视觉语言建模实现鲁棒的少样本OOD检测。该方法引入三项创新:空间特征增强以提升场景区分能力;双提示对齐机制,跨模态验证场景上下文与细粒度语义的一致性;以及置信度引导的自训练循环,动态挖掘伪标签扩充训练数据,无需人工标注。RS-OOD在多个遥感基准上持续优于现有方法,且仅需少量标注即可高效适配,验证了空间-语义融合的关键价值。

原文摘要 · Abstract (English)

Out-of-distribution (OOD) detection represents a critical challenge in remote sensing applications, where reliable identification of novel or anomalous patterns is essential for autonomous monitoring, disaster response, and environmental assessment. Despite remarkable progress in OOD detection for natural images, existing methods and benchmarks remain poorly suited to remote sensing imagery due to data scarcity, complex multi-scale scene structures, and pronounced distribution shifts. To this end, we propose RS-OOD, a novel framework that leverages remote sensing-specific vision-language modeling to enable robust few-shot OOD detection. Our approach introduces three key innovations: spatial feature enhancement that improved scene discrimination, a dual-prompt alignment mechanism that cross-verifies scene context against fine-grained semantics for spatial-semantic consistency, and a confidence-guided self-training loop that dynamically mines pseudo-labels to expand training data without manual annotation. RS-OOD consistently outperforms existing methods across multiple remote sensing benchmarks and enables efficient adaptation with minimal labeled data, demonstrating the critical value of spatial-semantic integration.

遥感图像异常检测视觉语言模型少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。