arXiv:2606.30030cs.CV2026-06

模仿鹰眼机制,让模糊图像恢复更准更真实

CogSENet: Blind Image Deblurring with Blur-Conditioned Semantic Routing and Explicit Frequency Fusion

论文配图:CogSENet: Blind Image Deblurring with Blur-Conditioned Semantic Routing and Explicit Frequency Fusion
图 1 · 摘自论文原文
  • 用语义路由动态重组特征,捕捉长距离依赖关系
  • 通过小波分解高低频信息,提升纹理结构恢复精度
  • 结合模糊场与语义先验,自适应处理不均匀模糊

盲图像去模糊需从复杂未知退化中恢复高保真细节与连贯结构。现有方法难以应对真实世界中空间变化的退化,且缺乏语义感知能力,无法可靠区分有效纹理与伪影。为此,我们提出 CogSENet,一种受鹰眼视觉系统启发的动态语义对齐重建框架。通过模拟鹰眼的主动扫视行为,设计语义驱动的状态空间模块(SDSSM),利用可微路由实现语义感知的标记重分组,支持提示条件下的长程依赖建模。为确保纹理与结构的物理可解释性恢复,提出双频融合模块(BFFB),通过小波变换将特征分解为高低频成分,模拟鹰眼视网膜的功能分化。最后,从模糊图像估计连续模糊场(CBF),并融合CLIP语义先验,调制深层潜在特征,模拟聚焦调节,实现非均匀模糊下的自适应修复。大量实验表明,CogSENet在视觉质量与结构保真度上优于当前最优方法,参数更少,同时在去雾、去雨、去噪任务中也表现优异。

原文摘要 · Abstract (English)

Blind image deblurring demands the recovery of high-fidelity details and coherent structures from complex, unknown degradations. Current blind image deblurring methods struggle with real-world, spatially varying degradations, and lack the semantic awareness necessary to reliably differentiate valid textures from artifacts. To bridge this gap, we propose CogSENet, a dynamic, semantic-aligned reconstruction framework inspired by the eagle's visual system. By mimicking the eagle's active saccadic scanning, we devise a Semantic-Driven State Space Module (SDSSM) with semantic-aware token regrouping via differentiable routing, enabling prompt-conditioned long-range dependency modeling. To ensure physically interpretable recovery of textures and structures, a BiFreqFusionBlock (BFFB) mirrors functional differentiation of the eagle's retina by decomposing features into high and low frequencies using wavelet transforms. Finally, we estimate a continuous Blur Field (CBF) from blur image and fuse it with CLIP semantic priors to modulate the deepest latent features, emulating focal adaptation and enabling adaptive restoration under spatially non-uniform blur. Extensive experiments demonstrate that CogSENetoutperforms state-of-the-art deblurring methods in both visual quality and structural fidelity with fewer parameters, while also performing favorably on dehazing, deraining, and denoising tasks.

图像去模糊语义路由多频融合鹰眼模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。