融合物理模型与胶囊网络,实现高效高质水下图像增强
Physics Informed Capsule Enhanced Variational AutoEncoder for Underwater Image Enhancement
- 双流架构:物理模型估计透射图与背景光,胶囊聚类提取实体特征
- 比最优方法提升0.5dB PSNR,计算量仅为三分之一
- 无需调参,适合对细节和语义结构要求高的水下视觉应用
我们提出一种新颖的双流架构,通过显式整合Jaffe-McGlamery物理模型与基于胶囊聚类的特征表示学习,实现当前最优的水下图像增强。该方法在专用物理估计流中同时估计透射图与空间变化的背景光,另一并行流则通过胶囊聚类提取实体级特征。这种物理引导的方法实现了无参数增强,既符合水下成像约束,又保留语义结构与细粒度细节。此外,我们设计了一种新优化目标,确保在多尺度空间频率下兼具物理一致性与感知质量。在六个挑战性基准上进行了广泛实验,结果表明,相比现有最优方法,本方法在保持仅1/3计算复杂度(FLOPs)的前提下,提升0.5dB PSNR;若对比同类计算预算的方法,则提升超过1dB PSNR。代码与数据将发布于https://github.com/iN1k1/。
原文摘要 · Abstract (English)
We present a novel dual-stream architecture that achieves state-of-the-art underwater image enhancement by explicitly integrating the Jaffe-McGlamery physical model with capsule clustering-based feature representation learning. Our method simultaneously estimates transmission maps and spatially-varying background light through a dedicated physics estimator while extracting entity-level features via capsule clustering in a parallel stream. This physics-guided approach enables parameter-free enhancement that respects underwater formation constraints while preserving semantic structures and fine-grained details. Our approach also features a novel optimization objective ensuring both physical adherence and perceptual quality across multiple spatial frequencies. To validate our approach, we conducted extensive experiments across six challenging benchmarks. Results demonstrate consistent improvements of $+0.5$dB PSNR over the best existing methods while requiring only one-third of their computational complexity (FLOPs), or alternatively, more than $+1$dB PSNR improvement when compared to methods with similar computational budgets. Code and data \textit{will} be available at https://github.com/iN1k1/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。