arXiv:2603.01140cs.CV2026-03

通过因果干预分离图像内容与噪声,提升去噪鲁棒性。

Teacher-Guided Causal Interventions for Image Denoising: Orthogonal Content-Noise Disentanglement in Vision Transformers

  • 在视觉变换器中施加结构化干预,显式分解内容与噪声表征。
  • 在多个基准上超越主流方法,单卡实现实时104.2帧速度。
  • 适合关注图像去噪泛化能力与真实场景表现的研究者。

传统图像去噪模型常错误学习环境因素与噪声模式之间的虚假关联。由于高频模糊性,它们难以可靠区分细微纹理与随机噪声,导致细节过度去除或残留噪声。本文重新审视去噪问题,认为仅依赖相关性拟合会混淆内在内容与外生噪声,直接降低分布偏移下的鲁棒性。为此提出教师引导的因果解耦网络(TCD-Net),在视觉变换器框架内通过结构化干预显式分解生成机制。具体包括:(1) 环境偏差调整(EBA)模块将特征投影至稳定、去中心化子空间,抑制全局环境偏差(去混淆);(2) 双分支解耦头引入正交性约束,强制内容与噪声表征严格分离,防止信息泄露;(3) 为解决结构模糊性,借助 Google 的推理引导生成模型 Nano Banana Pro 提供因果先验,有效将内容表征拉回自然图像流形。大量实验表明,TCD-Net 在多个基准上均优于主流方法,在保真度与效率方面表现优异,单张 RTX 5090 GPU 实现 104.2 FPS 实时速度。

原文摘要 · Abstract (English)

Conventional image denoising models often inadvertently learn spurious correlations between environmental factors and noise patterns. Moreover, due to high-frequency ambiguity, they struggle to reliably distinguish subtle textures from stochastic noise, resulting in over-removed details or residual noise artifacts. We therefore revisit denoising via causal intervention, arguing that purely correlational fitting entangles intrinsic content with extrinsic noise, which directly degrades robustness under distribution shifts. Motivated by this, we propose the Teacher-Guided Causal Disentanglement Network (TCD-Net), which explicitly decomposes the generative mechanism via structured interventions on feature spaces within a Vision Transformer framework. Specifically, our method integrates three key components: (1) An Environmental Bias Adjustment (EBA) module projects features into a stable, de-centered subspace to suppress global environmental bias (de-confounding). (2) A dual-branch disentanglement head employs an orthogonality constraint to force a strict separation between content and noise representations, preventing information leakage. (3) To resolve structural ambiguity, we leverage Nano Banana Pro, Google's reasoning-guided AI image generation model, to guide a causal prior, effectively pulling content representations back onto the natural-image manifold. Extensive experiments demonstrate that TCD-Net outperforms mainstream methods across multiple benchmarks in both fidelity and efficiency, achieving a real-time speed of 104.2 FPS on a single RTX 5090 GPU.

图像去噪因果推理视觉变换器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。