用脑科学方法发现大模型能自动捕捉人类视觉兴趣
Neuroscience-Inspired Analyses of Visual Interestingness in Multimodal Transformers

- 通过神经科学分析法解码模型内部表征中的视觉兴趣信号
- 中间视觉层已显现出可区分的兴趣特征,语言层更强化此模式
- 无需标注数据即形成结构化兴趣编码,适合研究注意力机制者
人类注意力是意识感知、记忆与决策的入口,但其在现代Transformer模型中的作用仍不明确。随着这些系统日益影响人们的观看、偏好与消费行为,关键问题在于它们是否编码了人类兴趣原理,还是仅利用大规模相关性。为回应这一挑战,本文基于社交平台Flickr的大规模人类参与数据构建的通用兴趣度(CI)评分,在多模态视觉-语言模型Qwen3-VL-8B中分析了视觉兴趣的内在表征。采用神经科学研究方法,分析了视觉与语言组件的内部表示。结果表明,最终层嵌入可线性解码出CI信息,说明其与人类视觉兴趣测量高度一致。降维与广义判别值(GDV)分析显示,与CI相关的隐藏表征在中间视觉变压器层出现,并随语言模型层推进而逐步增强可区分性。几何、探测器及稀疏自编码器方法提取的概念向量在高层趋于收敛,经表示相似性分析验证,表明视觉兴趣被无监督地稳健编码。未来工作将探索人脑动态与变压器架构间的共享计算原则,目标是揭示生物与人工系统中注意力与兴趣的组织机制。
原文摘要 · Abstract (English)
Human attention is the gateway to conscious perception, memory and decision-making. However, its role in modern transformer models remains largely unexplored. As these systems increasingly influence what people see, prefer and buy, the question arises as to whether they encode principles of human interest or merely exploit large-scale correlations. Addressing this issue is crucial for understanding cognition and ensuring the responsible use of AI in communication and marketing. In order to address this issue, the concept of visual interest was examined within the multimodal vision-language-model Qwen3-VL-8B, using a pre-defined Common Interestingness (CI) score derived from large-scale human engagement data on the photo-sharing platform Flickr. Here, we analyzed internal representations across vision and language components using methods from the neurosciences. Our analyses revealed that CI information is linearly decodable from final-layer embeddings, indicating that it is aligned with human-derived measures of visual interestingness. Dimensionality reduction and Generalized Discrimination Value (GDV) analyses demonstrate that CI-related hidden representations emerge in intermediate vision transformer layers and becomes progressively more distinguishable across language model layers. Concept vectors derived using geometric, probe, and Sparse Auto-Encoder based methods converge in higher layers, as confirmed by representational similarity analysis. This indicates a robust and structured encoding of visual interestingness without explicit supervision. Future work will seek to identify shared computational principles linking human brain dynamics and transformer architectures, with the ultimate goal of uncovering the organizing mechanisms that give rise to attention and interest in both biological and artificial systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。