arXiv:2604.07254cs.CVcs.LG2026-04

深度模型能准确预测人类真伪判断,但解释不一致,难说其推理逻辑。

Non-identifiability of Explanations from Model Behavior in Deep Networks of Image Authenticity Judgments

论文配图:Non-identifiability of Explanations from Model Behavior in Deep Networks of Image Authenticity Judgments
图 1 · 摘自论文原文
  • 用多种归因方法测试模型解释的稳定性,发现跨架构一致性差。
  • 模型预测准确率达噪声上限的80%,但解释结果差异大。
  • 适合关注模型可解释性局限的研究者,警惕过度解读归因图。

深度神经网络可预测人类对图像真伪的判断,但这并不意味着它们使用了与人类相似的信息或揭示了判断依据。以往研究依赖归因热图解释模型行为,但其有效性取决于鲁棒性。本文通过评估不同架构下预测人类真伪评分的模型是否产生一致的归因,检验了归因的稳健性。在多个冻结的预训练视觉模型上添加轻量回归头,并使用Grad-CAM、LIME和多尺度像素掩码生成归因图。部分架构预测表现良好,达到噪声上限的约80%。其中VGG模型主要依赖图像质量而非真伪特异性特征,削弱了归因相关性。其余模型在单个架构内归因图对随机种子变化较稳定,尤其在EfficientNetB3和Barlow Twins中更明显,且对被判断为更真实的图像一致性更高。然而,即使预测性能相似,跨架构间的归因一致性仍很弱。为此,我们采用模型集成提升人类真伪判断预测效果,并通过像素掩码实现图像级归因。结论表明:尽管深度网络能良好预测人类真伪判断,但无法提供可识别的解释。更广泛而言,成功行为模型的后验解释应视为认知机制的弱证据。

原文摘要 · Abstract (English)

Deep neural networks can predict human judgments, but this does not imply that they rely on human-like information or reveal the cues underlying those judgments. Prior work has addressed this issue using attribution heatmaps, but their explanatory value in itself depends on robustness. Here we tested the robustness of such explanations by evaluating whether models that predict human authenticity ratings also produce consistent explanations within and across architectures. We fit lightweight regression heads to multiple frozen pretrained vision models and generated attribution maps using Grad-CAM, LIME, and multiscale pixel masking. Several architectures predicted ratings well, reaching about 80% of the noise ceiling. VGG models achieved this by tracking image quality rather than authenticity-specific variance, limiting the relevance of their attributions. Among the remaining models, attribution maps were generally stable across random seeds within an architecture, especially for EfficientNetB3 and Barlow Twins, and consistency was higher for images judged as more authentic. Crucially, agreement in attribution across architectures was weak even when predictive performance was similar. To address this, we combined models in ensembles, which improved prediction of human authenticity judgments and enabled image-level attribution via pixel masking. We conclude that while deep networks can predict human authenticity judgments well, they do not produce identifiable explanations for those judgments. More broadly, our findings suggest that post hoc explanations from successful models of behavior should be treated as weak evidence for cognitive mechanism.

可解释性图像真伪模型归因深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。