用简单模型从猴子脑神经信号还原视觉,准确率达70%
Simple Models, Rich Representations: Visual Decoding from Primate Intracortical Neural Signals
- 用时序注意力+浅层MLP捕捉神经信号动态,提升解码精度
- 200毫秒脑活动可还原图像,顶1准确率最高达70%
- 适合脑机接口与语义神经解码研究者参考
理解神经活动如何产生感知是神经科学的核心挑战。我们针对灵长类动物高密度皮层内记录数据,使用THINGS腹侧流尖峰数据集,系统评估了模型架构、训练目标和数据规模对解码性能的影响。结果表明,解码准确率主要由对神经信号时序动态的建模决定,而非模型复杂度。一个结合时序注意力与浅层MLP的简单模型,在图像检索任务中达到最高70%的顶1准确率,优于线性基线及循环、卷积方法。规模分析显示,随着输入维度和数据集规模增加,性能提升呈现可预测的递减趋势。基于此,我们设计了一种模块化生成解码流程,结合低分辨率潜在空间重建与语义条件扩散模型,仅需200毫秒脑活动即可生成合理图像。该框架为脑机接口与语义神经解码提供了原则性指导。
原文摘要 · Abstract (English)
Understanding how neural activity gives rise to perception is a central challenge in neuroscience. We address the problem of decoding visual information from high-density intracortical recordings in primates, using the THINGS Ventral Stream Spiking Dataset. We systematically evaluate the effects of model architecture, training objectives, and data scaling on decoding performance. Results show that decoding accuracy is mainly driven by modeling temporal dynamics in neural signals, rather than architectural complexity. A simple model combining temporal attention with a shallow MLP achieves up to 70% top-1 image retrieval accuracy, outperforming linear baselines as well as recurrent and convolutional approaches. Scaling analyses reveal predictable diminishing returns with increasing input dimensionality and dataset size. Building on these findings, we design a modular generative decoding pipeline that combines low-resolution latent reconstruction with semantically conditioned diffusion, generating plausible images from 200 ms of brain activity. This framework provides principles for brain-computer interfaces and semantic neural decoding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。