无需训练,动态校准遥感图像语义分割中的视觉与文本一致性。
Seeking Consensus: Geometric-Semantic On-the-Fly Recalibration for Open-Vocabulary Remote Sensing Semantic Segmentation

- 通过多视角几何一致性和文本自适应校准,实现在线协同重校准。
- 在8个遥感数据集上均实现性能提升,显著缓解前景激活不足问题。
- 适用于各类零样本语义分割模型,部署简单且通用性强。
遥感图像中的开放词汇语义分割(OVSS)是一项有前景的任务,利用文本描述识别未定义的地表类别。尽管已有显著进展,现有方法通常采用静态推理范式,忽视了不同场景间分布差异,导致地表类型语义模糊和前景激活不全。为此,我们提出「寻求共识」(SeeCo),一个即插即用的框架,可对无训练的OVSS模型进行实时校准。SeeCo通过双共识机制实现动态重校准:基于多视图一致观测的几何共识学习(GCL)与基于文本描述自适应校准的语义共识学习(SCL),两者通过在线共识注入器(OCI)协同作用,有效缓解语义偏差与激活不足。该方法无需额外训练,在推理阶段针对每个独特场景重校准视觉-语义对齐。在8个遥感图像OVSS基准上的实验表明,该方法持续取得性能提升,验证了其有效性与普适性。
原文摘要 · Abstract (English)
Open-vocabulary semantic segmentation (OVSS) in remote sensing images is a promising task that employs textual descriptions for identifying undefined land cover categories. Despite notable advances, existing methods typically employ a static inference paradigm, overlooking the distinct distribution of each scene, resulting in semantic ambiguity in diverse land covers and incomplete foreground activation. Motivated by this, we propose Seeking Consensus, termed SeeCo, a plug-and-play framework to boost the performance of training-free OVSS models in remote sensing images, which recalibrates arbitrary OVSS models on-the-fly by seeking dual consensus: geometric consensus learning (GCL) through multi-view consistent observations and semantic consensus learning (SCL) via textual description adaptive calibration, which assists collaborative recalibration of visual and textual semantics. The two consensus are injected via an online consensus injector (OCI), effectively alleviating the under-activation and semantic bias. SeeCo requires no specific training process, yet recalibrates semantic-geometric alignment for each unique scene during inference. Extensive experiments on eight remote sensing OVSS benchmarks show consistent gains, proving its effectiveness and universality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。