arXiv:2504.21214cs.CLcs.AI2025-04被引 11

用120小时脑电数据预训练大模型,提升静默语音解码准确率。

Pretraining Large Brain Language Model for Active BCI: Silent Speech

  • 提出FSTP自回归预训练方法,捕捉脑电信号时空依赖
  • 跨会话场景下语义分类达47.0%,词级分类39.6%,显著优于基线
  • 首个面向主动脑机接口的静默语音大模型,适合神经工程研究者

本文探索主动脑机接口中静默语音解码问题,该技术相较传统应用更具自然性和灵活性。研究收集了来自12名受试者的超120小时脑电(EEG)数据,涵盖24个常用英文单词,用于语言模型预训练与解码。受大型模型自监督预训练成功启发,本文提出大型脑语言模型(LBLM)以解码静默语音。为预训练LBLM,提出未来频时预测(FSTP)预训练范式,通过时间与频率域的自回归建模,从无标签脑电数据中学习有效表征。不同于现有主要采用掩码重建的预训练方法,FSTP能同时捕捉信号的时序与谱依赖性。预训练后,在词级与语义级分类等下游任务上进行微调。大量实验表明,LBLM在多个任务上显著优于全监督及预训练基线模型。例如在困难的跨会话设置下,语义级分类准确率达47.0%,词级分类达39.6%,分别领先基线5.4%和7.3%。本研究推动了主动脑机接口中的静默语音解码进展,提出了创新的脑电语言模型预训练方案,并提供了新的基础研究数据集。

原文摘要 · Abstract (English)

This paper explores silent speech decoding in active brain-computer interface (BCI) systems, which offer more natural and flexible communication than traditional BCI applications. We collected a new silent speech dataset of over 120 hours of electroencephalogram (EEG) recordings from 12 subjects, capturing 24 commonly used English words for language model pretraining and decoding. Following the recent success of pretraining large models with self-supervised paradigms to enhance EEG classification performance, we propose Large Brain Language Model (LBLM) pretrained to decode silent speech for active BCI. To pretrain LBLM, we propose Future Spectro-Temporal Prediction (FSTP) pretraining paradigm to learn effective representations from unlabeled EEG data. Unlike existing EEG pretraining methods that mainly follow a masked-reconstruction paradigm, our proposed FSTP method employs autoregressive modeling in temporal and frequency domains to capture both temporal and spectral dependencies from EEG signals. After pretraining, we finetune our LBLM on downstream tasks, including word-level and semantic-level classification. Extensive experiments demonstrate significant performance gains of the LBLM over fully-supervised and pretrained baseline models. For instance, in the difficult cross-session setting, our model achieves 47.0\% accuracy on semantic-level classification and 39.6\% in word-level classification, outperforming baseline methods by 5.4\% and 7.3\%, respectively. Our research advances silent speech decoding in active BCI systems, offering an innovative solution for EEG language model pretraining and a new dataset for fundamental research.

脑机接口静默语音大模型脑电分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。