用二维码结构测试解释方法是否真懂图像结构
CAMBench-QR : A Structure-Aware Benchmark for Post-Hoc Explanations with QR Understanding
- 用二维码几何特征设计可控制的评测数据集
- 发现多数方法在关键结构上注意力不足,背景干扰多
- 适合评估模型解释是否真正理解图像结构
视觉解释常看似合理却缺乏结构忠实性。我们提出CAMBench-QR,一个基于二维码标准几何结构(定位图案、时序线、模块网格)的结构感知评测基准,用于检验CAM方法是否在必要子结构上分配显著性且避免背景干扰。该基准通过精确掩码和可控扭曲生成二维码与非二维码数据,报告结构感知指标(定位/时序质量比、背景泄漏、覆盖率AUC、距结构距离)以及因果遮蔽、插入/删除忠实性、鲁棒性和延迟等。我们在零样本和最后层微调两种实用场景下,对代表性高效CAM方法(LayerCAM、EigenGrad-CAM、XGrad-CAM)进行评测。该基准、度量与训练方案提供了一个简单、可复现的结构感知评估基准。因此我们建议将CAMBench-QR作为检验视觉解释是否真正结构感知的试纸。
原文摘要 · Abstract (English)
Visual explanations are often plausible but not structurally faithful. We introduce CAMBench-QR, a structure-aware benchmark that leverages the canonical geometry of QR codes (finder patterns, timing lines, module grid) to test whether CAM methods place saliency on requisite substructures while avoiding background. CAMBench-QR synthesizes QR/non-QR data with exact masks and controlled distortions, and reports structure-aware metrics (Finder/Timing Mass Ratios, Background Leakage, coverage AUCs, Distance-to-Structure) alongside causal occlusion, insertion/deletion faithfulness, robustness, and latency. We benchmark representative, efficient CAMs (LayerCAM, EigenGrad-CAM, XGrad-CAM) under two practical regimes of zero-shot and last-block fine-tuning. The benchmark, metrics, and training recipes provide a simple, reproducible yardstick for structure-aware evaluation of visual explanations. Hence we propose that CAMBENCH-QR can be used as a litmus test of whether visual explanations are truly structure-aware.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。