arXiv:2607.27268cs.SD2026-07

测试脑电大模型能否用于语音解码,发现现有预训练无效

Does EEG Foundation Models Transfer to Speech? A Benchmark on Overt and Imagined Speech Decoding

  • 用统一流程对比大模型与卷积网络在语音解码上的表现
  • 在两个数据集上,大模型性能不如16K参数的小卷积网
  • 提示需构建专门的语音脑电预训练模型

基于数千小时数据预训练的脑电大模型在运动想象、癫痫检测、睡眠分期和情绪识别任务中表现出显著优势,但在最具挑战性的非侵入式脑机接口应用——语音解码领域尚未验证。本文首次系统性地将脑电大模型(LaBraM、EEGMamba)与三种强基线模型(EEGNet、ShallowFBCSPNet、EEGConformer)在两个语料库上进行对比:UGR-MINDVOICE(口语与默念伊比利亚西班牙语)和BCI Competition 2020 Track 3(想象语音)。所有模型在统一预处理与微调协议下比较。结果表明,大规模脑电预训练在语音任务上未展现出一致优势,其性能甚至不及仅含16,000参数的卷积神经网络,说明当前通用脑电预训练尚未有效迁移至语音生成任务,亟需构建面向语音的专用脑电基础模型。

原文摘要 · Abstract (English)

EEG foundation models pretrained on thousands of hours have shown large gains over task-specific networks for motor imagery, seizure detection, sleep staging, and emotion recognition, but their transfer to speech decoding-arguably the most demanding non-invasive BCI application-remains untested. We present the first systematic benchmark of EEG foundation models against strong convolutional baselines for speech decoding, using two corpora: UGR-MINDVOICE (overt and covert Iberian Spanish) and BCI Competition 2020 Track 3 (imagined speech). We compare two foundation models (LaBraM, EEGMamba) against three established baselines (EEGNet, ShallowFBCSPNet, EEGConformer) under a unified preprocessing and fine-tuning protocol. Large-scale EEG pretraining yields no consistent advantage over a 16K-parameter CNN on speech tasks, indicating that current general-purpose EEG pretraining does not yet transfer to speech production and motivating speech-specific foundation models.

脑电解码语音生成大模型预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。