用脑电图解码无声阅读中的词汇语义,效果显著且可扩展。
Decoding silent reading from non-invasive EEG

- 以无声阅读为代理任务,用对比学习从脑电中提取词汇信息。
- 在约24万词数据上实现高于随机的词级解码,罕见词也有效。
- 结果表明解码能力受限于数据量而非模型上限,适合神经工程研究者。
非侵入式解码内在言语面临根本性数据难题:无法收集大脑活动与自发内心独白的配对数据集。现有替代范式(提示重复和事后报告生成式内言)获取慢、时间锁不准,且无法验证受试者配合度。因此,我们采用无声阅读作为可扩展的代理任务,探究对比解码器能从中提取多少词汇与语义信息。本研究对单个高密度采样受试者进行了393次运行(约49小时),使用19通道干电极EEG记录了约24万词的快速序列视觉呈现数据,文本为连续叙事内容,每轮试验字体随机化以部分消除词身份与低层视觉形式的相关性。采用卷积EEG编码器,可选接因果Transformer,通过类似CLIP的对比目标,将短时EEG窗口与大语言模型中对应词的隐状态嵌入对齐。解码性能以词组级前10检索相对于置换基线评估,结果稳定高于随机水平,涵盖中频与罕见词,并随训练数据量呈对数线性增长,未见饱和迹象。移除枕叶与后颞叶电极使词级增益下降约三分之一,但上下文追踪能力保持不变。控制分析区分了词级解码、叙事上下文追踪以及由Transformer位置嵌入引入的非神经位置先验。结果表明,无声阅读期间可恢复开放词汇的词级信息,且解码能力受数据限制而非饱和。
原文摘要 · Abstract (English)
Non-invasive decoding of inner speech faces a fundamental data problem: a corpus pairing brain activity with a person's spontaneous inner monologue cannot be collected, and the available proxy paradigms (cued repetitive and retrospectively reported generative inner speech) are slow to acquire, poorly time-locked, and subject compliance is unverifiable. We therefore treat silent reading as a scalable proxy task and ask how much lexical and semantic information a contrastive decoder can extract from it. We report an open-vocabulary analysis of approximately 240,000 word presentations recorded from a single densely-sampled participant across 393 runs (ca. 49 h) of 19-channel dry-electrode EEG. Words from continuous narrative text were presented in rapid serial visual presentation, with typography randomised on every trial to partially decorrelate word identity from low-level visual form. A convolutional EEG encoder, optionally followed by a causal transformer, was trained with a CLIP-style contrastive objective to align short EEG windows with hidden-state embeddings of the presented word taken from a large language model. Decoding, evaluated as word-grouped top-10 retrieval against permutation baselines, was reliably above chance, extended to mid-frequency and rare words, and scaled log-linearly with training-data volume with no sign of saturation. Removing occipital and posterior-temporal electrodes reduced the word-level gain by roughly one third but left context tracking unchanged. Control analyses separate word-level decoding from narrative context tracking and from a non-neural positional prior introduced by the transformer's positional embedding. These results establish that open-vocabulary word-level information is recoverable from EEG during silent reading, and that decoding is data-limited rather than saturated.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。