提升脑电解码语音的跨人泛化能力,降低训练成本。
Cross-Subject Generalization in Decoding Perceived Speech from Non-Invasive Brain Recordings

- 分两阶段训练:先用对比学习学共性特征,再针对个体微调。
- 在三个数据集上性能超越基线6.8%以上,最高达15.8%。
- 适合需要跨被试通用的脑机接口研究者使用。
从非侵入式脑记录中解码感知语音近年来受到广泛关注,但跨被试解码面临泛化能力差、缺乏提取共性信息机制等问题,导致训练成本高、性能不佳。为此,本文提出跨被试感知语音解码(CPSD)框架,包含源模型预训练与个性化专精两个阶段:在预训练阶段采用对比学习捕捉多源被试间的共享表示;在专精阶段通过提取源模型中的一致成分并结合目标被试数据微调,实现个性化适配。此外,引入基于位置编码的空间注意力(PESA)模块,将MEG/EEG数据映射至标准化参考空间,增强跨被试一致性,促进模型训练。在涵盖不同模态和语言的三个感知语音神经数据集上评估,CPSD在Armeni 2022、PKUEEG 2025、Broderick 2018数据集上的Top-10准确率分别优于基线6.8%、15.4%、15.8%。进一步分析验证了该方法的有效性、高效性与鲁棒性。
原文摘要 · Abstract (English)
Decoding perceived speech from non-invasive brain recordings has garnered significant attention in recent years due to its wide range of potential applications. However, existing methods face considerable challenges in cross-subject decoding, primarily due to limited generalizability and the absence of explicit mechanisms for extracting subject-consistent information. These limitations result in high training costs and suboptimal decoding performance. To address these challenges, we propose an innovative Cross-Subject Perceived Speech Decoding (CPSD) framework, which comprises two training stages: source model pre-training and personal specialization. In the source model pre-training stage, contrastive learning is employed to capture shared representations across multiple source subjects. Subsequently, personal specialization initializes the model for the target subject by extracting consistent components from the source model and fine-tuning it using target subject data. Additionally, we introduce the Positional Encoding-based Spatial Attention (PESA) module, which remaps MEG/EEG data into a standardized reference space, thereby enhancing cross-subject consistency and facilitating model training. We evaluate the proposed CPSD framework on three perceived speech neural datasets encompassing different modalities and languages. The results demonstrate that our framework outperforms baseline methods by more than 6.8%, 15.4%, and 15.8% in Top-10 accuracy on the Armeni 2022, PKUEEG 2025, and Broderick 2018 datasets, respectively. Further analyses confirm the effectiveness, efficiency, and robustness of the proposed approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。