用脑电肌电融合解码中文四声,跨人通用性强且只需少量电极。
CAT-Net: A Cross-Attention Tone Network for Cross-Subject EEG-EMG Fusion Tone Decoding
- 通过交叉注意力融合脑电与肌电信号,模拟言语产生中神经与肌肉协作机制。
- 在仅20个脑电通道和5个肌电通道下,静默说话准确率达88.08%,跨人平均85.10%。
- 适合开发轻量级、可跨人使用的中文语音脑机接口系统,尤其助残应用。
脑机接口(BCI)语音解码为言语障碍者提供了新工具。本研究提出一种跨被试多模态BCI解码框架,融合脑电(EEG)与肌电(EMG)信号,在发声与静默说话条件下对四种汉语声调进行分类。受言语生成中神经-肌肉协同机制启发,模型采用时空特征提取分支与交叉注意力融合机制,实现模态间有效交互,并引入领域对抗训练以提升跨被试泛化能力。实验采集了10名参与者共4,800次EEG和4,800次EMG数据,仅使用20个EEG通道和5个EMG通道,验证了极简通道解码的可行性。尽管模块轻量化,模型在所有条件下均优于现有基线,发声条件平均准确率达87.83%,静默说话达88.08%;跨被试评估中,发声与静默分别达到83.27%与85.10%。消融实验证明各组件有效性。结果表明,使用极少通道实现声调级解码可行且具跨被试泛化潜力,推动实用化BCI发展。
原文摘要 · Abstract (English)
Brain-computer interface (BCI) speech decoding has emerged as a promising tool for assisting individuals with speech impairments. In this context, the integration of electroencephalography (EEG) and electromyography (EMG) signals offers strong potential for enhancing decoding performance. Mandarin tone classification presents particular challenges, as tonal variations convey distinct meanings even when phonemes remain identical. In this study, we propose a novel cross-subject multimodal BCI decoding framework that fuses EEG and EMG signals to classify four Mandarin tones under both audible and silent speech conditions. Inspired by the cooperative mechanisms of neural and muscular systems in speech production, our neural decoding architecture combines spatial-temporal feature extraction branches with a cross-attention fusion mechanism, enabling informative interaction between modalities. We further incorporate domain-adversarial training to improve cross-subject generalization. We collected 4,800 EEG trials and 4,800 EMG trials from 10 participants using only twenty EEG and five EMG channels, demonstrating the feasibility of minimal-channel decoding. Despite employing lightweight modules, our model outperforms state-of-the-art baselines across all conditions, achieving average classification accuracies of 87.83% for audible speech and 88.08% for silent speech. In cross-subject evaluations, it still maintains strong performance with accuracies of 83.27% and 85.10% for audible and silent speech, respectively. We further conduct ablation studies to validate the effectiveness of each component. Our findings suggest that tone-level decoding with minimal EEG-EMG channels is feasible and potentially generalizable across subjects, contributing to the development of practical BCI applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。