arXiv:2608.05122cs.CV2026-08

用脑科学方法分析视觉Transformer如何学会识别方向,揭示其早期具备生物特性特征。

IRIS: A Visual Cortex-Inspired Framework for Analyzing Orientation Selectivity in Vision Transformers

论文配图:IRIS: A Visual Cortex-Inspired Framework for Analyzing Orientation Selectivity in Vision Transformers
图 1 · 摘自论文原文
  • 借鉴神经科学指标,量化ViT中方向选择性的编码机制。
  • 训练初期即出现方向选择性单元,中层最多,深层逐渐消失转向语义编码。
  • 可指导冻结层数以提升下游任务泛化能力,适合模型可解释性研究者。

视觉变换器(ViTs)已成为图像编码的主流方法,但其对低级特征(如方向选择性)的编码机制仍不清晰,因其缺乏归纳偏置,处理方式为全局而非局部。生物视觉系统则通过小范围视觉区域信息整合形成方向选择性等通用表征,广泛用于多条神经通路。本研究引入一系列神经科学启发的指标:表示相似性得分(RSS)、方向激活得分(ORS)和方向调谐带宽,系统分析了方向选择性在ViT中的演化过程。结果表明:(1) 训练范式是决定方向选择性的最强因素,不同规模模型在相对深度相近时达到峰值;(2) 许多神经元在训练早期就具备方向选择性,中层随时间招募更多此类单元,而深层则丧失选择性并拓宽调谐带宽,转向语义编码;(3) 所提指标可作为优化微调策略的机制性启发,帮助确定最佳解冻层数。该框架为追踪ViT训练过程中生物相关特征提供新路径,深化对表征编码与跨任务泛化的理解。

原文摘要 · Abstract (English)

Vision transformers (ViTs) have become the de facto standard for image encoding across many perception tasks. Despite their empirical success, it remains mechanistically unclear how they encode low-level features, given their lack of inductive biases: ViTs process information globally rather than relying on local structure. Biological visual systems, in contrast, build low-level features, such as orientation selectivity in the primary visual cortex, by combining information from small, localized regions of the visual field. These features are general-purpose representations, shared and required across multiple specialized neural pathways, unlike higher-level, task-specific semantic features. This raises the question if such biologically-grounded features arise in ViTs. In this work, we systematically study how orientation selectivity emerges in ViTs by introducing a suite of neuroscience-inspired metrics: representational similarity score (RSS), orientation recruitment score (ORS), and orientation tuning bandwidth to quantify how orientation is encoded in representational geometry and as a function of model depth. Through extensive analysis, we find that: (1) the training paradigm is the strongest determinant of orientation selectivity, with models sharing an objective, peaking at comparable relative depths regardless of scale (2) many units are orientation-selective early in training, with early-to-middle layers recruiting more such units over time, while deeper layers lose selectivity and broaden their tuning toward semantic encoding and (3) our metrics offer a mechanistic heuristic for how many layers to unfreeze for best downstream generalization. Our framework presents a way to track biologically-grounded features during ViT training, probes how desired properties are encoded in transformer representations, and builds a systematic understanding of how ViTs generalize across tasks.

视觉Transformer方向选择性可解释性神经科学启发

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。