解决遥感图像变化描述中的视角、尺度模糊问题,提升描述准确性。
STAND: Semantic Anchoring Constraint with Dual-Granularity Disambiguation for Remote Sensing Image Change Captioning

- 引入语义锚定约束,净化时序特征表示
- 双粒度消歧模块提升大范围视角与小物体识别精度
- 结合语言先验知识,缓解解码阶段的知识模糊
遥感图像变化描述(RSICC)旨在刻画两张遥感图像间的差异。现有方法虽尝试视频建模,但普遍忽视视点、尺度及先验知识带来的固有模糊性,缺乏对编码器的有效约束。本文提出STAND,一种基于语义锚定与双粒度消歧的遥感图像变化描述方法,以逐步消除这些模糊性。首先,通过可解释的约束正则化时序表示,建立可靠特征基础;在此基础上,双粒度消歧模块通过宏观全局上下文聚合缓解视点混淆,并利用微观频域聚焦注意力增强小目标尺度感知;最终,语义概念锚定模块借助语言类别先验,在解码阶段应对知识模糊。大量实验验证了STAND的优越性及其在消除模糊性方面的有效性。
原文摘要 · Abstract (English)
Remote sensing image change captioning (RSICC) aims to describe the difference between two remote sensing images. While recent methods have explored video modeling, they largely overlook the inherent ambiguities in viewpoint, scale, and prior knowledge, lacking effective constraints on the encoder. In this paper, we present STAND, a Semantic Anchoring Constraint with Dual-Granularity Disambiguation for RSICC, to progressively resolve these ambiguities. Specifically, to establish a reliable feature foundation, we first introduce an interpretable constraint to regularize temporal representations. Operating on these purified features, a dual-granularity disambiguation module resolves spatial uncertainties by coupling macro-level global context aggregation for viewpoint confusion with micro-level frequency-refocused attention for small-object scale enhancement. Ultimately, to translate these visually disambiguated features into precise text, a semantic concept anchoring module leverages language categorical priors to tackle knowledge ambiguity during decoding. Extensive experiments verify the superiority of STAND and its effectiveness in addressing ambiguities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。