arXiv:2605.00025q-bio.NCcs.CL2026-05中稿 · ICMI 2026

通过解耦机制发现大脑中互补的言语神经信号,提升脑机接口翻译准确率。

MoDAl: Self-Supervised Neural Modality Discovery via Decorrelation for Speech Neuroprosthesis

  • 利用对比对齐与解相关双目标,在共享空间中挖掘不同脑区的互补信号。
  • 在脑到文本基准上将词错误率从26.3%降至21.6%,提升来自区域44的信号价值。
  • 揭示了区域44编码句法与结构特征,符合布罗卡区的神经语言学功能。

言语神经假体系统通过解析无声音输出时的大脑活动来还原意图言语,为言语障碍者提供沟通恢复路径。现有方法主要依赖运动皮层,忽略如布罗卡区部分的区域44等可能携带互补语言信息的区域。我们提出MoDAl(模态解耦与对齐)框架,通过共享投影空间中两个目标的协同作用,发现互补神经模态。对比损失将多个并行脑区编码器与预训练大语言模型(LLM)的文本嵌入对齐,而解相关损失防止编码器趋于冗余表示。我们证明这两个目标存在建设性张力:对比对齐会诱导模态同化,需解相关机制加以抑制,才能发现多样化的神经语言模态。在Brain-to-Text Benchmark '24上,相比此前最优端到端方法,MoDAl将词错误率(WER)从26.3%降低至21.6%,其中新增区域44信号带来的增益完全源于解相关机制。对发现模态的分析显示,接收区域44输入的编码器捕捉句子长度、语法语态、疑问词等句法结构特征,与布罗卡区的神经语言学认知一致。

原文摘要 · Abstract (English)

Speech neuroprosthesis systems decode intended speech from neural activity in the absence of audible output, offering a path to restoring communication for individuals with speech-impairing conditions. Current approaches decode predominantly from motor cortical areas, discarding others -- such as area 44, part of Broca's area -- that may encode complementary linguistic information. We introduce MoDAl (Modality Decorrelation and Alignment), a framework that discovers complementary neural modalities through the interplay of two objectives in a shared projection space. A contrastive loss aligns each of several parallel brain encoders with the text embeddings of a pretrained large language model (LLM), while a decorrelation loss prevents the encoders from coalescing to duplicative representations. We prove that these objectives are in productive tension: Contrastive alignment induces transitive modality coalescence, which decorrelation must counteract for the framework to discover diverse neurolinguistic modalities. On the Brain-to-Text Benchmark '24, MoDAl reduces word error rate (WER) from 26.3% to 21.6% compared to the previous best end-to-end method, with the gain from incorporating previously discarded area 44 signals arising entirely from the decorrelation mechanism. Analysis of the discovered modalities reveals functional specialization: Encoders receiving area 44 input capture structural and syntactic properties (sentence length, grammatical voice, wh-words), consistent with the neurolinguistic understanding of Broca's area.

脑机接口神经解码多模态解耦学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。