arXiv:2608.09995eess.IVcs.CV2026-08

用结构先验提升相机图像去马赛克与去噪的稳定性。

Structural Guidance for Unified Joint Demosaicing and Denoising

论文配图:Structural Guidance for Unified Joint Demosaicing and Denoising
图 1 · 摘自论文原文
  • 引入结构推理分支,从稀疏伪RGB中提取互补结构线索。
  • 在多种滤色阵列和噪声水平下优于当前最优方法。
  • 适合需要高保真图像重建的相机系统研发人员。

联合去马赛克与去噪是相机图像信号处理中的基础步骤,但因不同类似Bayer的色彩滤波阵列(CFAs)与传感器噪声共同破坏颜色采样与图像内容,仍具挑战性。现有统一恢复网络虽显式建模CFA几何结构,但仍主要依赖像素级监督,导致边缘、重复纹理和莫尔条纹区域易出现结构退化,因局部证据不可靠。我们发现,这部分源于缺乏超越像素级重建的显式结构引导。为此,提出一种结构引导的统一恢复框架,将预训练结构知识注入面向CFA的图像恢复。模型接收包含原始马赛克、CFA掩码和噪声水平图的五通道统一观测。一个SwinIR恢复分支在CFA条件调制下重建像素细节,同时并行的结构推理分支从稀疏伪RGB观测中提取互补结构线索。为弥合稀疏噪声传感器数据与结构编码器预训练自然图像域之间的巨大领域差异,我们在残差融合前引入轻量级可训练适配器。共享解码器联合预测恢复后的RGB图像与辅助干净马赛克,在图像与传感器双域提供监督。跨多种CFA模式与噪声水平的大量实验表明,该方法持续优于最先进的统一及特定于CFA的方法,证明适配的结构先验能增强鲁棒的相机图像恢复能力。源代码与数据集见附录。

原文摘要 · Abstract (English)

Joint demosaicing and denoising is a fundamental step in camera image signal processing, yet remains challenging because different Bayer-like color filter arrays (CFAs) and sensor noise jointly corrupt both color sampling and image content. Existing unified restoration networks explicitly model CFA geometry but are still driven primarily by pixel-level supervision, making them prone to structural degradation around edges, repetitive textures, and moiré patterns where local evidence is unreliable. We attribute this limitation partly to the absence of explicit structural guidance beyond pixel-level reconstruction supervision. Motivated by this observation, we propose a structural-guided unified restoration framework that injects pretrained structural knowledge into CFA-aware image restoration. Our model receives a unified five-channel observation consisting of the raw mosaic, CFA masks, and a noise-level map. A SwinIR restoration branch reconstructs pixel details under CFA-conditioned modulation, while a parallel structural reasoning branch extracts complementary structural cues from a sparse pseudo-RGB observation. To bridge the substantial domain gap between sparse noisy sensor data and the natural-image pretraining domain of the structural encoder, we introduce a lightweight trainable adapter before residually fusing structural and restoration features. A shared decoder jointly predicts the restored RGB image and an auxiliary clean mosaic, providing supervision in both image and sensor domains. Extensive experiments across multiple CFA patterns and noise levels demonstrate consistent improvements over state-of-the-art unified and CFA-specific methods, indicating that adapted structural priors can enhance robust camera image restoration. The source codes and dataset are provided in the supplementary material.

图像恢复去马赛克结构先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。