用脑电数据直接检索对应句子,靠对比学习对齐大脑与语言
sEEG-based Encoding for Sentence Retrieval: A Contrastive Learning Approach to Brain-Language Alignment
- 用对比学习将脑电信号映射到预训练语言模型的句子空间
- 在有限数据下实现脑活动到句子的准确检索
- 无需微调文本模型,适合脑-语言研究者使用
通过有意义的潜在表示来解释神经活动,仍是神经科学与人工智能交叉领域中一个复杂且不断演进的挑战。我们探究了多模态基础模型在对侵入性脑记录与自然语言进行对齐方面的潜力。提出SSENSE,一种对比学习框架,将单个受试者的立体脑电图(sEEG)信号投影到冻结的CLIP模型的句子嵌入空间,实现仅凭脑活动进行句子级检索。SSENSE在频谱表示的sEEG上训练神经编码器,使用InfoNCE损失,不微调文本编码器。我们在一个自然场景电影观看数据集上评估方法,该数据集包含时间对齐的sEEG与口语转录文本。尽管数据有限,SSENSE仍取得良好结果,表明通用语言表示可作为神经解码的有效先验。
原文摘要 · Abstract (English)
Interpreting neural activity through meaningful latent representations remains a complex and evolving challenge at the intersection of neuroscience and artificial intelligence. We investigate the potential of multimodal foundation models to align invasive brain recordings with natural language. We present SSENSE, a contrastive learning framework that projects single-subject stereo-electroencephalography (sEEG) signals into the sentence embedding space of a frozen CLIP model, enabling sentence-level retrieval directly from brain activity. SSENSE trains a neural encoder on spectral representations of sEEG using InfoNCE loss, without fine-tuning the text encoder. We evaluate our method on time-aligned sEEG and spoken transcripts from a naturalistic movie-watching dataset. Despite limited data, SSENSE achieves promising results, demonstrating that general-purpose language representations can serve as effective priors for neural decoding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。