提出CHEM框架,量化图像重建中的幻觉区域并分析成因。
CHEM: Estimating and Understanding Hallucinations in Deep Learning for Image Processing
- 用小波与剪切波表示定位幻觉区域,结合分位数回归实现无分布假设的评估。
- 在天文图像去卷积和自然图像超分辨率任务中,验证了不同模型的幻觉水平差异。
- 揭示U型网络易产生幻觉的理论根源,为安全关键场景提供可解释性工具。
基于深度学习的方法在图像重建任务中取得显著进展,但可能生成不真实的人工伪影或幻觉,影响安全关键场景下的分析。本文提出一种量化与表征图像重建模型中幻觉伪影的框架——共形幻觉估计度量(CHEM)。该方法利用小波与剪切波表示,在图像特征层面定位幻觉高风险区域,并通过共形化分位数回归以无分布假设方式评估幻觉程度。我们提供了理论分析,刻画了CHEM对幻觉伪影的敏感性及其与均方误差的关系。基于逼近论视角,进一步探究了广泛使用的U型网络为何倾向于产生幻觉预测。我们在天文图像去卷积任务(使用CANDELS数据集)和自然图像超分辨率任务(使用DIV2K数据集)上评估了该方法的有效性,涵盖U-Net、SwinUNet、Learnlets、DRUNet、Unfolded DRS、RAM和DPS等模型。
原文摘要 · Abstract (English)
Deep learning-based methods have recently achieved significant success in image reconstruction problems. However, challenges have emerged, as these methods may generate unrealistic artifacts or hallucinations, which can interfere with analysis in safety-critical scenarios. This paper introduces a framework for quantifying and characterizing hallucinated artifacts in image reconstruction models. The proposed method, termed the Conformal Hallucination Estimation Metric (CHEM), enables the identification of hallucination-prone regions in model predictions. It leverages wavelet and shearlet representations to localize such regions at the level of image features, and uses conformalized quantile regression to assess hallucination levels in a distribution-free manner. A theoretical analysis is provided, characterizing the sensitivity of CHEM to hallucinated artifacts and its relationship to the mean squared error. Building on these insights and adopting a viewpoint grounded in approximation theory, we investigate why U-shaped networks, widely used architectures for image reconstruction, tend to hallucination-prone predictions. We assess the effectiveness of the proposed approach on astronomical image deconvolution using the CANDELS dataset with architectures such as U-Net, SwinUNet, and Learnlets, and on natural image super-resolution using the DIV2K dataset with models such as DRUNet, Unfolded DRS, RAM, and DPS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。