用语义扫描与三令牌提示提升高阶RWKV的图像融合效果
Multigrain-aware Semantic Prototype Scanning and Tri-Token Prompt Learning Embraced High-Order RWKV for Pan-Sharpening

- 基于语义聚类的多粒度原型扫描,实现上下文感知的特征重排
- 三令牌机制增强语义先验,抑制噪声与伪影,提升细节保留
- 无参可逆Q-Shift操作高效注入高频信息,保持特征完整性
本文提出一种面向全色锐化任务的多粒度语义原型扫描范式,基于高阶RWKV架构与源自语义聚类的三令牌提示机制。方法包含三部分:1)多粒度语义原型扫描。尽管RWKV具备线性复杂度优势,但传统双向光栅扫描仍缺乏语义感知且易受位置偏差影响。为此,引入基于局部敏感哈希的语义驱动扫描策略,将语义相关区域分组并构建多粒度语义原型,实现上下文感知的标记重排与更连贯的全局交互。2)三令牌提示学习。设计包含全局令牌、聚类生成原型令牌和可学习寄存器令牌的三令牌机制,全局与原型令牌为RWKV建模提供互补语义先验,寄存器令牌则有助于抑制噪声及伪影明显的中间表示。3)可逆Q-Shift。为对抗空间细节损失,在值路径上应用中心差分卷积注入高频信息,并引入可逆多尺度Q-Shift操作,实现高效无损特征变换,避免参数密集的感受野扩展。实验结果表明该方法具有显著优越性。
原文摘要 · Abstract (English)
In this work, we propose a Multigrain-aware Semantic Prototype Scanning paradigm for pan-sharpening, built upon a high-order RWKV architecture and a tri-token prompting mechanism derived from semantic clustering. Specifically, our method contains three key components: 1) Multigrain-aware Semantic Prototype Scanning. Although RWKV offers a efficient linear-complexity alternative to Transformers, its conventional bidirectional raster scanning is still semantic-agnostic and prone to positional bias. To address this issue, we introduce a semantic-driven scanning strategy that leverages locality-sensitive hashing to group semantically related regions and construct multi-grain semantic prototypes, enabling context-aware token reordering and more coherent global interaction. 2) Tri-token Prompt Learning. We design a tri-token prompting mechanism consisting of a global token, cluster-derived prototype tokens, and a learnable register token. The global and prototype tokens provide complementary semantic priors for RWKV modeling, while the register token helps suppress noisy and artifact-prone intermediate representations. 3) Invertible Q-Shift. To counteract spatial details, we apply center difference convolution on the value pathway to inject high-frequency information, and introduce an invertible multi-scale Q-shift operation for efficient and lossless feature transformation without parameter-heavy receptive field expansion. Experimental results demonstrate the superiority of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。