arXiv:2605.20337cs.CV2026-05被引 1

发现视觉大模型特征不如监督模型可解释,且与性能无关

Capability $\neq$ Interpretability: Human Interpretability of Vision Foundation Models

论文配图:Capability $\neq$ Interpretability: Human Interpretability of Vision Foundation Models
图 1 · 摘自论文原文
  • 用人类感知实验测量模型特征的定位和命名能力
  • 大模型特征可解释性普遍低于监督模型,且与任务性能无关
  • 局部激活和粗粒度语义对齐是提升可解释性的关键

视觉基础模型的特征有多可解释?随着这些模型进入高风险应用,这一问题愈发紧迫。现有方法无法可靠回答。本文提出一个基于心理物理学的评估框架,通过两个互补实验:(1)定位性——观察者能否预测特征在新图像上的激活位置;(2)命名性——能否准确描述特征所代表的内容。使用稀疏自编码器提取特征,并以随机基线为锚点进行评分,使不同模型可在同一尺度比较。对六种视觉变压器(两种监督ViT,四种基础模型:DINOv2、DINOv3、CLIP、SigLIP)进行测试,共收集超过15,000次行为响应,分析377名通过质量检查的参与者提交的13,400条有效数据。结果表明,基础模型的特征可解释性始终低于其监督对应模型,且该差距并非能力权衡——可解释性与任何下游任务表现均无相关性。真正相关的因素是特征激活的局部性以及与人类粗粒度语义结构的对齐程度。具有集中激活且反映世界基本分类结构的模型产生更易理解的特征,而细粒度感知对齐则无效。两项实验排名高度一致,共享相同预测因子,确立了可解释性作为表示质量独立可测维度的地位,且所有测试的基础模型均不及此前的监督基线。仅靠能力不足以弥补差距,局部性与粗粒度对齐才是关键。

原文摘要 · Abstract (English)

How interpretable are the features of leading vision models? The question is increasingly pressing as these models move from research benchmarks into high-stakes deployments, yet existing methods cannot answer it reliably. We close this gap with a framework for measuring and comparing the human interpretability of vision models, built around two complementary psychophysics protocols: (1) localizability -- can an observer predict where a feature fires on a novel image? -- and (2) nameability -- can an observer accurately describe what the feature represents? Features are recovered via sparse autoencoders, and a chance-anchored scoring function places every model on a common scale. Applying the framework to six vision transformers -- two supervised ViTs and four foundation models (DINOv2, DINOv3, CLIP, SigLIP) -- we collected more than $15{,}000$ behavioral responses, analyzing the $13{,}400$ responses from the $377$ participants who passed our pre-specified quality checks. Foundation models are consistently *less* interpretable than their supervised counterparts, and the gap is not a capability tradeoff: interpretability does not correlate with downstream task performance on any benchmark we examine. What does correlate is the locality of a feature's activations and coarse-grained semantic alignment with humans -- models with focal activations and representations that reflect the world's broad categorical structure produce more interpretable features, whereas fine-grained perceptual alignment does not. The two protocols yield strongly correlated rankings and share the same predictors, establishing interpretability as an independent, measurable dimension of representation quality -- and, surprisingly, one on which every foundation model we tested falls below the supervised baselines that came before. Capability alone cannot close that gap; locality and coarse-grained alignment can.

可解释性视觉模型大模型心理物理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。