纠正图像分解评估漏洞,提出更可靠的场景级划分标准
The frame-level leakage trap: rethinking evaluation protocols for intrinsic image decomposition, with source-separable uncertainty as a case study

- 采用场景级划分替代帧级划分,避免数据泄露
- 帧级划分使测试指标虚高1.6~2.0 dB,长期训练可超10 dB
- 首次实现光源分离的不确定性建模,提升图像重建质量
MPI Sintel数据集上基于学习的固有图像分解评估协议存在不一致问题。先前研究采用帧级划分,导致同一场景中空间相似的帧同时出现在训练与测试集中,造成数据泄露。本文首次量化该泄漏效应:在三种架构下,帧级划分使测试R_PSNR虚增1.6至2.0 dB(p<0.01,配对t检验,3次随机种子),且该效应与模型无关。三重梯度分析(随机/时间/场景)表明该差距呈连续分布,扩展训练后虚增超过10 dB。本文倡导以场景级划分为社区标准,并提供六种代表性模型在此协议下的基准性能。作为案例研究,在修正协议下,提出物理启发的分解模型 I = R ⊙ S + N,配备光源分离的三路异方差不确定性头。实证验证通道特异性:非朗伯反射不确定性通道与非朗伯残差误差的交叉相关系数达r=0.67,是纹理通道的4倍以上。进一步证明下游效用:剔除75%最高不确定性像素后,保留像素的重建均方误差降低77%,而随机剔除无改善。该特性在分布外真实照片上依然成立。报告了包含频域分解、跨任务监督、证据学习、对比损失与测试时自适应的复杂变体的负面结果。本方法达到15.98±0.41 dB R_PSNR,仅需五分之一成本即可接近5成员深度集成(差0.8 dB),且具备唯一光源分离不确定性能力。
原文摘要 · Abstract (English)
Evaluation protocols for learned intrinsic image decomposition on MPI Sintel have been inconsistent. Several prior works split the dataset by frames, which allows spatially similar frames of the same scene to appear in both train and test partitions. We quantify this leakage effect for the first time, across three architectures: a frame-level split inflates test R_PSNR by 1.6 to 2.0 dB (p less than 0.01 for all three, paired t-test across 3 seeds) relative to a scene-level split, confirming an architecture-independent protocol effect. A three-point gradient (random/temporal/scene) shows the gap is continuous, and under extended training the frame-level inflation exceeds 10 dB. We advocate scene-level splits as the community standard and provide reference numbers for six representative models under this protocol. As a case study within the corrected protocol, we present a physics-informed decomposition I = R composed with S + N with a source-separable three-way heteroscedastic uncertainty head. We empirically verify channel specialization: the non-Lambertian uncertainty channel shows r = 0.67 cross-correlation with non-Lambertian residual error, more than 4 times the texture channel's correlation. We further demonstrate downstream utility: filtering out the 75% highest-uncertainty pixels reduces reconstruction MSE by 77% on retained pixels, whereas random filtering produces no improvement. The specialization also holds on out-of-distribution real photographs. We report negative results for a more elaborate variant combining frequency decomposition, cross-task supervision, evidential learning, contrastive loss, and test-time adaptation. Our method reaches 15.98 plus or minus 0.41 dB R_PSNR, within 0.8 dB of a 5-member Deep Ensemble at one-fifth the cost, with the unique capability of source-separated uncertainty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。