提出新评估框架,更可靠地诊断模型对形状与纹理的依赖程度。
On the Reliability of Cue Conflict and Beyond
- 基于明确定义的形状和纹理构造可辨识的提示对
- 在全标签空间测量线索敏感性,避免偏差误判
- 解决旧方法中线索有效性与识别度混淆问题
理解神经网络如何依赖视觉线索,有助于揭示其决策机制。现有基于风格化的提示冲突基准在探测形状-纹理偏好时存在不稳定性,因风格化难以可靠生成感知有效且可分离的线索,且比率型偏差指标会掩盖绝对敏感性,限制类别评估则可能因忽略完整决策空间而扭曲模型表现。这些因素导致偏好被误判为线索有效性、平衡性或可识别性伪影。为此,我们提出REFINED-BIAS——一个集成数据集与评估框架,通过显式定义形状与纹理,构建人类与模型均可识别的平衡线索对,并采用基于排序的度量方式,在全标签空间上量化线索特异性敏感性,实现更公平的跨模型比较。在多种训练策略与架构下,REFINED-BIAS显著提升偏差诊断的准确性,获得更清晰的实证结论,解决了以往评估无法可靠区分的矛盾现象。
原文摘要 · Abstract (English)
Understanding how neural networks rely on visual cues offers a human-interpretable view of their internal decision processes. The cue-conflict benchmark has been influential in probing shape-texture preference and in motivating the insight that stronger, human-like shape bias is often associated with improved in-domain performance. However, we find that the current stylization-based instantiation can yield unstable and ambiguous bias estimates. Specifically, stylization may not reliably instantiate perceptually valid and separable cues nor control their relative informativeness, ratio-based bias can obscure absolute cue sensitivity, and restricting evaluation to preselected classes can distort model predictions by ignoring the full decision space. Together, these factors can confound preference with cue validity, cue balance, and recognizability artifacts. We introduce REFINED-BIAS, an integrated dataset and evaluation framework for reliable and interpretable shape-texture bias diagnosis. REFINED-BIAS constructs balanced, human- and model- recognizable cue pairs using explicit definitions of shape and texture, and measures cue-specific sensitivity over the full label space via a ranking-based metric, enabling fairer cross-model comparisons. Across diverse training regimes and architectures, REFINED-BIAS enables fairer cross-model comparison, more faithful diagnosis of shape and texture biases, and clearer empirical conclusions, resolving inconsistencies that prior cue-conflict evaluations could not reliably disambiguate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。