arXiv:2608.03147cs.CV2026-08中稿 · ECCV被引 1

解决遥感图像指代分割的定位漂移与语义偏差问题

CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation

论文配图:CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation
图 1 · 摘自论文原文
  • 通过级联蒸馏将SAM结构先验注入VLM中间层,增强定位精度
  • 引入双约束对比学习,有效抵抗主导对象语义干扰
  • 在复杂空间描述下仍保持精准分割,适合高精度遥感分析

遥感图像指代分割(RRSIS)得益于视觉语言模型(VLM)与通用分割模型(SAM)的融合,取得显著进展。然而,当前方法过度依赖预训练能力,存在两大根本缺陷:(1)架构弱耦合,单向信息流导致依赖粗粒度的VLM提示,浪费SAM的像素级结构引导,引发定位漂移;(2)以对象为中心的语义偏差,模型过度关注主导对象语义,忽视对空间推理的敏感性。为此,本文提出CROSS,一种紧密集成的RRSIS新范式。首先,设计语言引导级联蒸馏(LGCD),将SAM的几何亲和性作为软正则化注入VLM中间层,注入密集结构先验以优化定位。其次,提出视角-空间对比学习(PSCL),通过挖掘掩码过滤后的误导性干扰项和空间-语言反事实样本作为难负样本,显式打破语义捷径,强制逻辑一致性。在多个RRSIS基准测试中,CROSS达到领先性能,并在严重空间描述扰动下仍保持精确分割,成为鲁棒的新范式。

原文摘要 · Abstract (English)

Referring Remote Sensing Image Segmentation (RRSIS) has achieved significant progress through the integration of VLMs and the Segment Anything Model (SAM). However, this progress largely relies on strong pre-trained capabilities, while leaving two fundamental limitations insufficiently addressed: (1) Architectural Weak-Coupling, where the unidirectional flow forces reliance on coarse VLM prompts and wastes SAM's pixel-level structural guidance, causing localization drift; and (2) Object-Centric Semantic Bias, where models overemphasize dominant object semantics while remaining insensitive to spatial reasoning crucial for RRSIS. Motivated by these observations, we propose CROSS, a tightly integrated paradigm for RRSIS. First, we introduce Linguistic-Guided Cascaded Distillation (LGCD) to bridge the architectural gap, which distills SAM's geometric affinities as soft regularizers into VLM intermediate layers, injecting dense structural priors to refine localization. Second, Perspective-Spatial Contrastive Learning (PSCL) imposes cross-anchored constraints by mining mask-filtered deceptive distractors and spatial-linguistic counterfactuals as hard negatives, explicitly shattering semantic shortcuts to enforce genuine logical consistency. Extensive experiments on RRSIS benchmarks demonstrate that CROSS achieves state-of-the-art performance and maintains precise localization even under severe spatial description perturbations, standing as a robust new paradigm for RRSIS.

遥感分割视觉语言模型结构引导对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。