arXiv:2605.23144cs.CV2026-05

用结构化属性替代传统标签,提升遥感目标检测的细粒度识别与泛化能力

SLIP-RS: Structured-Attribute Language-Image Pre-Training for Remote Sensing Object Detection

论文配图:SLIP-RS: Structured-Attribute Language-Image Pre-Training for Remote Sensing Object Detection
图 1 · 摘自论文原文
  • 将开放类别映射为有限物理属性空间,通过组合增强学习解耦视觉逻辑
  • 在1500万属性标注数据上训练,实现跨域检测性能新突破
  • 适合遥感图像分析、细粒度目标检测研究者使用

现有遥感目标检测的语言-图像预训练受限于单体标签学习,依赖黑箱数据枚举开放集类别以获取细粒度表征,导致与遥感领域数据稀缺特性不兼容。为此,我们提出SLIP-RS,建立结构化属性解耦范式,将开放类别空间映射为有限且具物理意义的属性空间,通过显式结构逻辑解锁细粒度判别能力。该范式由两大技术支柱实现:(1) 结构化属性对比学习,通过组合属性增强强制学习解耦的内在视觉逻辑;(2) 保形属性可靠性引擎,利用保形预测理论从噪声源中严格提炼高保真监督信号,生成包含超过1500万属性标注的RS-Attribute-15M数据集。大量实验表明,SLIP-RS在细粒度检测和跨域泛化方面达到前所未有的性能,验证了结构化属性作为遥感基础的重要价值。

原文摘要 · Abstract (English)

Existing language-image pre-training for remote sensing object detection is constrained by Monolithic Label Learning, which relies on exhaustively enumerating open-set categories via black-box data to acquire fine-grained representations, creating a dependency incompatible with the domain's inherent data scarcity. To transcend this bottleneck, we propose SLIP-RS, establishing a Structured-Attribute Decoupling Paradigm that maps the open-ended category space into a finite, physically meaningful attribute space, unlocking fine-grained discriminability via explicit structural logic. This paradigm is realized via two technical pillars: (1) Structured-Attribute Contrastive Learning, which enforces the learning of decoupled intrinsic visual logic via combinatorial attribute augmentation; and (2) Conformal Attribute Reliability Engine, which leverages conformal prediction theory to rigorously distill high-fidelity supervision from noisy sources, yielding RS-Attribute-15M, the largest dataset with over 15 million attribute annotations. Extensive experiments demonstrate that SLIP-RS establishes unprecedented performance in fine-grained detection and cross-domain generalization, validating structured attributes as a vital foundation for remote sensing. Code: https://github.com/facias914/SLIP-RS.

遥感检测属性学习对比学习数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。