arXiv:2505.02705eess.IVcs.CV2025-05IJCAI被引 5

用上下文引导的注意力机制,高效修复真实场景中的复杂噪声。

Multi-View Learning with Context-Guided Receptance for Image Denoising

  • 引入上下文感知的令牌迁移机制,捕捉局部空间依赖关系。
  • 在多个真实图像去噪数据集上超越现有方法,推理速度提升40%。
  • 适合需要快速高精度去噪的自动驾驶与摄影应用。

图像去噪在摄影和自动驾驶等低层视觉任务中至关重要。现有方法难以区分真实场景中的复杂噪声模式,且因依赖Transformer模型导致计算开销大。本文提出上下文引导的接收权重键值(Context-guided Receptance Weighted Key-Value, \\(M\\)模型,结合增强的多视角特征融合与高效序列建模。引入上下文引导的令牌迁移(CTS)范式,有效捕获局部空间依赖,提升对真实噪声分布的建模能力。此外,设计频域特征提取的频率混合(FMix)模块,分离高频谱中的噪声,并通过多视角学习与空间表示融合。为提升计算效率,采用双向WKV(BiWKV)机制,在保持线性复杂度的同时实现全像素序列交互,突破因果选择限制。模型在多个真实世界图像去噪数据集上验证,定量结果优于现有最先进方法,推理时间最多减少40%。定性结果进一步证明模型在各类场景下恢复细节的能力。

原文摘要 · Abstract (English)

Image denoising is essential in low-level vision applications such as photography and automated driving. Existing methods struggle with distinguishing complex noise patterns in real-world scenes and consume significant computational resources due to reliance on Transformer-based models. In this work, the Context-guided Receptance Weighted Key-Value (\M) model is proposed, combining enhanced multi-view feature integration with efficient sequence modeling. Our approach introduces the Context-guided Token Shift (CTS) paradigm, which effectively captures local spatial dependencies and enhance the model's ability to model real-world noise distributions. Additionally, the Frequency Mix (FMix) module extracting frequency-domain features is designed to isolate noise in high-frequency spectra, and is integrated with spatial representations through a multi-view learning process. To improve computational efficiency, the Bidirectional WKV (BiWKV) mechanism is adopted, enabling full pixel-sequence interaction with linear complexity while overcoming the causal selection constraints. The model is validated on multiple real-world image denoising datasets, outperforming the existing state-of-the-art methods quantitatively and reducing inference time up to 40\%. Qualitative results further demonstrate the ability of our model to restore fine details in various scenes.

图像去噪多视角学习高效建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。