arXiv:2510.02182q-bio.NCcs.CV2025-10被引 1

用扩散模型揭示视觉皮层神经群的语义选择性,直接看到它们如何编码物体特征。

Uncovering Semantic Selectivity of Latent Groups in Higher Visual Cortex with Mutual Information-Guided Diffusion

  • 通过变分自编码器分离神经潜空间,再用互信息引导扩散模型生成特定视觉语义
  • 在猕猴下颞叶皮层数据中验证,不同神经群对物体姿态、类别间变换等有明确语义选择
  • 适合研究神经编码机制或想理解大脑如何组织视觉信息的计算神经科学家

理解高级视觉区域神经群体如何编码以物体为中心的视觉信息,仍是计算神经科学的核心挑战。以往研究多通过人工神经网络与视觉皮层表征对齐来间接推断,缺乏对神经群体结构的深入洞察;解码方法虽能量化语义特征,却未揭示其内在组织方式。本研究提出MIG-Vis方法,利用扩散模型的生成能力,可视化并验证神经潜空间子群所编码的视觉-语义属性。该方法首先通过变分自编码器从神经群体中推断出分组解耦的神经潜空间;随后设计互信息(MI)引导的扩散合成流程,生成每组潜变量对应的特定视觉-语义特征。我们在两只猕猴的下颞叶(IT)皮层多会话神经放电数据上验证了该方法。结果表明,所识别的神经潜空间组对多种视觉特征具有清晰的语义选择性,包括物体姿态、跨类别变换及类内内容差异。这些发现为高级视觉皮层中结构化的语义表征提供了直接且可解释的证据,深化了我们对其编码原理的理解。

原文摘要 · Abstract (English)

Understanding how neural populations in higher visual areas encode object-centered visual information remains a central challenge in computational neuroscience. Prior works have investigated representational alignment between artificial neural networks and the visual cortex. Nevertheless, these findings are indirect and offer limited insights to the structure of neural populations themselves. Similarly, decoding-based methods have quantified semantic features from neural populations but have not uncovered their underlying organizations. This leaves open a scientific question: "how feature-specific visual information is distributed across neural populations in higher visual areas, and whether it is organized into structured, semantically meaningful subspaces." To tackle this problem, we present MIG-Vis, a method that leverages the generative power of diffusion models to visualize and validate the visual-semantic attributes encoded in neural latent subspaces. Our method first uses a variational autoencoder to infer a group-wise disentangled neural latent subspace from neural populations. Subsequently, we propose a mutual information (MI)-guided diffusion synthesis procedure to visualize the specific visual-semantic features encoded by each latent group. We validate MIG-Vis on multi-session neural spiking datasets from the inferior temporal (IT) cortex of two macaques. The synthesized results demonstrate that our method identifies neural latent groups with clear semantic selectivity to diverse visual features, including object pose, inter-category transformations, and intra-class content. These findings provide direct, interpretable evidence of structured semantic representation in the higher visual cortex and advance our understanding of its encoding principles.

神经编码扩散模型视觉皮层语义表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。