发现低层视觉模型泛化失败主因是复杂背景导致的捷径学习。
Revisiting the Generalization Problem of Low-level Vision Models Through the Lens of Image Deraining
- 通过平衡内容与退化复杂度,引导模型关注真实图像重建。
- 引入预训练生成模型先验,约束网络在高质量图像流形上优化。
- 在去雨、去噪、去模糊任务中验证方法有效性,适用于鲁棒性研究者。
低层视觉模型在未见退化上的泛化能力仍是核心挑战。本文以去雨为例,系统实验揭示:泛化失败并非源于网络容量不足,而是由图像内容与退化模式之间的相对复杂度引发的“捷径学习”现象。当背景内容过复杂时,模型倾向于优先拟合更简单的退化特征以最小化训练损失,从而忽略底层图像分布。为此,提出两种原则性策略:(1) 平衡训练数据中背景与退化复杂度,引导模型聚焦内容重建;(2) 利用预训练生成模型中的强内容先验,物理性约束网络位于高质量图像流形上。在图像去雨、去噪和去模糊任务上的广泛实验验证了理论洞察。本工作提供了可解释性的视角与提升低层视觉模型鲁棒性与泛化能力的方法论。
原文摘要 · Abstract (English)
Generalization to unseen degradations remains a fundamental challenge for low-level vision models. This paper aims to investigate the underlying mechanism of this failure, using image deraining as a primary case study due to its well-defined and decoupled structure. Through systematic experiments, we reveal that generalization issues are not primarily caused by limited network capacity, but rather by a ``shortcut learning'' phenomenon driven by the relative complexity between image content and degradation patterns. We find that when background content is excessively complex, networks preferentially overfit the simpler degradation characteristics to minimize training loss, thereby failing to learn the underlying image distribution. To address this, we propose two principled strategies: (1) balancing the complexity of training data (backgrounds vs. degradations) to redirect the network's focus toward content reconstruction, and (2) leveraging strong content priors from pre-trained generative models to physically constrain the network onto a high-quality image manifold. Extensive experiments on image deraining, denoising, and deblurring validate our theoretical insights. Our work provides an interpretability-driven perspective and a principled methodology for improving the robustness and generalization of low-level vision models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。