用智能代理替代像素匹配,提升遥感图像少样本分割精度
AgMTR: Agent Mining Transformer for Few-shot Segmentation in Remote Sensing

- 通过自适应挖掘局部感知代理构建代理级语义关联
- 在iSAID数据集上达到当前最优性能,显著降低误分割率
- 适用于遥感与自然图像场景,泛化能力强
少样本分割(FSS)旨在仅用少量标注样本(支持图像)对查询图像中的目标进行分割。以往方法依赖支持-查询像素对之间的相似性构建像素级语义关联,但在具有极端类内差异和复杂背景的遥感场景中,此类关联易导致大量错配,造成前景与背景语义模糊。为此,本文提出新型代理挖掘变压器(AgMTR),自适应地挖掘一组具备局部上下文信息的代理,构建代理级语义关联。相比像素级语义,代理具有更广的感受野,使不同查询像素可选择性聚合不同代理的细粒度局部语义,从而增强查询图像中前景与背景的语义清晰度。具体地,提出代理学习编码器(ALE),建立最优传输方案,将不同代理分配至不同局部区域以聚合支持图像语义;进一步设计代理聚合解码器(AAD)与语义对齐解码器(SAD),分别从无标签数据源和查询图像自身挖掘有价值类别特定语义,突破支持集限制。在遥感基准数据集iSAID上的大量实验表明,所提方法达到当前最优性能。令人惊讶的是,该方法在更常见的自然图像场景(PASCAL-5i 和 COCO-20i)中也表现优异。
原文摘要 · Abstract (English)
Few-shot Segmentation (FSS) aims to segment the interested objects in the query image with just a handful of labeled samples (i.e., support images). Previous schemes would leverage the similarity between support-query pixel pairs to construct the pixel-level semantic correlation. However, in remote sensing scenarios with extreme intra-class variations and cluttered backgrounds, such pixel-level correlations may produce tremendous mismatches, resulting in semantic ambiguity between the query foreground (FG) and background (BG) pixels. To tackle this problem, we propose a novel Agent Mining Transformer (AgMTR), which adaptively mines a set of local-aware agents to construct agent-level semantic correlation. Compared with pixel-level semantics, the given agents are equipped with local-contextual information and possess a broader receptive field. At this point, different query pixels can selectively aggregate the fine-grained local semantics of different agents, thereby enhancing the semantic clarity between query FG and BG pixels. Concretely, the Agent Learning Encoder (ALE) is first proposed to erect the optimal transport plan that arranges different agents to aggregate support semantics under different local regions. Then, for further optimizing the agents, the Agent Aggregation Decoder (AAD) and the Semantic Alignment Decoder (SAD) are constructed to break through the limited support set for mining valuable class-specific semantics from unlabeled data sources and the query image itself, respectively. Extensive experiments on the remote sensing benchmark iSAID indicate that the proposed method achieves state-of-the-art performance. Surprisingly, our method remains quite competitive when extended to more common natural scenarios, i.e., PASCAL-5i and COCO-20i.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。