arXiv:2505.14074cs.HCcs.SD2025-05中稿 · presentation at In…被引 1

用大模型嵌入重建说话时脑电活动,效果极佳。

Recreating Neural Activity During Speech Production with Language and Speech Model Embeddings

  • 用语言和语音大模型的嵌入表示高阶语音特征
  • 所有参与者重建相关系数达0.79至0.99
  • 适合脑机接口与神经编码研究者参考

理解神经活动如何编码言语与语言生成,是神经科学与人工智能的核心挑战。本研究探讨大规模自监督语言与语音模型的嵌入是否能有效重建说话过程中记录的高伽马神经活动(高阶皮层处理的关键指标)。我们利用在语言与声学数据上预训练的深度学习模型嵌入,表征高层语音特征,并将其映射到高伽马信号上。分析这些嵌入对脑活动时空动态的保留程度。通过皮尔逊相关系数与信号重建质量评估,将重建信号与高伽马真实信号对比。结果表明,在所有受试者中,均可有效使用语言与语音模型嵌入重建高伽马活动,相关系数范围为0.79至0.99。

原文摘要 · Abstract (English)

Understanding how neural activity encodes speech and language production is a fundamental challenge in neuroscience and artificial intelligence. This study investigates whether embeddings from large-scale, self-supervised language and speech models can effectively reconstruct high-gamma neural activity characteristics, key indicators of cortical processing, recorded during speech production. We leverage pre-trained embeddings from deep learning models trained on linguistic and acoustic data to represent high-level speech features and map them onto these high-gamma signals. We analyze the extent to which these embeddings preserve the spatio-temporal dynamics of brain activity. Reconstructed neural signals are evaluated against high-gamma ground-truth activity using correlation metrics and signal reconstruction quality assessments. The results indicate that high-gamma activity can be effectively reconstructed using large language and speech model embeddings in all study participants, generating Pearson's correlation coefficients ranging from 0.79 to 0.99.

神经编码大模型嵌入脑机接口

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。