arXiv:2605.16468cs.CVcs.AI2026-05

用可解释模型揭示人脑视觉皮层对图像特征的精细选择性

Mechanistically Interpretable Neural Encoding Reveals Fine-Grained Functional Selectivity in Human Visual Cortex

论文配图:Mechanistically Interpretable Neural Encoding Reveals Fine-Grained Functional Selectivity in Human Visual Cortex
图 1 · 摘自论文原文
  • 通过语言对齐的图像表征,定位每个脑区像素对图像特征的响应机制
  • 预测的特征能生成与原图相似的激活模式,且优于随机或低置信度控制组
  • 可对图像特征进行反事实编辑,验证神经响应的因果关系,适合神经科学和AI交叉研究者

理解人类视觉的核心目标之一是揭示驱动神经元活动的视觉特征。现有方法多采用人工神经网络作为编码模型预测大脑对自然图像的反应,揭示了类别选择性区域被激活的内容,但这些方法大多是相关性的,将编码器视为黑箱,无法明确具体图像特征如何驱动每个体素(voxel)的响应。本文提出机械可解释神经编码框架(MINE),通过机制可解释工具打开这一黑箱,定位自然图像中驱动毫米级(体素级)活动的特征。MINE使用语言对齐的图像表征预测每个体素的响应,并生成语义可解释的激活关键特征描述。进一步,将这些每图像特征泛化为每体素的功能谱。为验证每图像描述的有效性,我们证明其足以生成激发匹配原始图像响应的图像,且准确率高于随机或低置信度控制组。此外,反事实地插入或移除预测特征会按预期方向改变激活水平,提供因果证据。基于每体素激活谱的反事实编辑产生更强的激活变化,表明该谱真实反映了体素的选择性。最后,我们将MINE应用于已知的类别选择性脑区,不仅复现了其已知的类别偏好,还揭示了各区域内独特的体素结构。整体结果表明,机制可解释性为发现并因果验证神经功能的细粒度假设提供了有效路径。

原文摘要 · Abstract (English)

A central goal in understanding human vision is to uncover the visual features that drive neuronal activity. A growing body of work has used artificial neural networks as encoding models to predict cortical responses to natural images, revealing the visual content that activates category-selective regions. However, existing approaches are largely correlational and treat the encoder as a black box, leaving open which image features drive each voxel's response. We introduce Mechanistically Interpretable Neural Encoding (MINE), a framework that opens this black box by applying mechanistic-interpretability tools to localize the features within natural images that drive millimeter-scale (voxel-level) activity. MINE predicts each voxel's response using language-aligned image representations, and produces semantically interpretable descriptions of the features critical for the voxel's activation. We further generalize these per-image features into per-voxel functional profiles. To validate the per-image descriptions, we show they are sufficient to generate images that elicit voxel responses matching the responses to the original images, more accurately than images generated from random or low-attribution controls. Moreover, counterfactually inserting or removing the predicted features from images shifts activation in the expected direction, providing causal evidence. Counterfactual editing guided by the per-voxel activation profiles produces even stronger activation shifts, indicating that the profiles faithfully capture each voxel's selectivity. Finally, we apply MINE to well-studied category-selective brain regions, showing it recovers their known categorical preferences while revealing fine-grained unique voxel structure within each region. Overall, our results establish mechanistic interpretability as a path to discover and causally validate fine-grained hypotheses about neural function.

神经编码可解释性视觉皮层因果推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。