arXiv:2511.21740cs.CLcs.AI2025-11被引 2

用统一模型直接把脑电转化成文字,比现有方法更准更快。

A cross-species neural foundation model for end-to-end speech decoding

  • 用跨任务跨物种的预训练神经编码器,统一处理说话意图和想象中的语音信号。
  • 在端到端模式下,将词错误率从24.69%降到10.22%,刷新纪录。
  • 适合脑机接口研究者和想用小模型提升语音解码性能的人。

语音脑机接口(BCI)旨在通过解析神经活动恢复瘫痪者的语言交流能力。现有系统多采用分步框架,先解码音素再用n-gram语言模型拼接句子,无法联合优化。本文提出端到端的脑电到文本(BIT)框架,使用单一可微神经网络直接生成连贯语句。核心是跨任务、跨物种预训练的神经编码器,其表征可迁移至实际发声与想象发声。在带n-gram语言模型的分步设置中,该编码器在Brain-to-Text '24和'25基准上达到新SOTA。与音频大语言模型(LLMs)端到端集成,并通过对比学习实现跨模态对齐后,词错误率(WER)从24.69%降至10.22%。值得注意的是,小规模音频LLMs显著提升端到端解码性能。此外,比特模型使实际发声与想象发声的嵌入对齐,支持跨任务泛化。整体上,该方法推动了大规模多样化神经数据的整合,为实现无缝、可微优化的端到端解码框架铺平道路。

原文摘要 · Abstract (English)

Speech brain-computer interfaces (BCIs) aim to restore communication for people with paralysis by translating neural activity into text. Most systems use cascaded frameworks that decode phonemes before assembling sentences with an n-gram language model (LM), preventing joint optimization of all stages simultaneously. Here, we introduce an end-to-end BraIn-to-Text (BIT) framework that translates neural activity into coherent sentences using a single differentiable neural network. Central to our approach is a cross-task, cross-species pretrained neural encoder, whose representations transfer to both attempted and imagined speech. In a cascaded setting with an n-gram LM, the pretrained encoder establishes a new state-of-the-art (SOTA) on the Brain-to-Text '24 and '25 benchmarks. Integrated end-to-end with audio large language models (LLMs) and trained with contrastive learning for cross-modal alignment, BIT reduces the word error rate (WER) of the prior end-to-end method from 24.69% to 10.22%. Notably, we find that small-scale audio LLMs markedly improve end-to-end decoding. Beyond record-setting performance, BIT aligns attempted and imagined speech embeddings to enable cross-task generalization. Altogether, our approach advances the integration of large, diverse neural datasets, paving the way for an end-to-end decoding framework that supports seamless, differentiable optimization.

脑机接口语音生成端到端大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。