用异构电极布局实现无声语音解码,提升患者识别准确率
A Silent Speech Decoding System from EEG and EMG with Heterogenous Electrode Configurations
- 采用多任务学习处理不同电极配置的脑电/肌电数据
- 健康人词分类准确率达95.3%,患者达54.5%,显著优于单人模型
- 适用于语言迁移和残障人士,推动实用化无声语音系统
无声语音解码通过脑电图/肌电图(EEG/EMG)实现无声言语识别,可提升言语障碍者的交流能力。但数据采集困难且实验设置各异,难以构建大规模同质数据集。本研究提出能处理异构电极布局的神经网络,在大规模EEG/EMG数据集上通过多任务训练,实现优异的无声语音解码性能。在健康受试者中达到95.3%的词分类准确率,在一名言语障碍患者中达54.5%,显著超越仅基于单人数据训练的模型(70.1%和13.2%)。此外,模型在跨语言校准任务中也表现更优。结果表明,该方法具备发展为实用无声语音系统的可行性,尤其对言语障碍患者具有重要意义。
原文摘要 · Abstract (English)
Silent speech decoding, which performs unvocalized human speech recognition from electroencephalography/electromyography (EEG/EMG), increases accessibility for speech-impaired humans. However, data collection is difficult and performed using varying experimental setups, making it nontrivial to collect a large, homogeneous dataset. In this study we introduce neural networks that can handle EEG/EMG with heterogeneous electrode placements and show strong performance in silent speech decoding via multi-task training on large-scale EEG/EMG datasets. We achieve improved word classification accuracy in both healthy participants (95.3%), and a speech-impaired patient (54.5%), substantially outperforming models trained on single-subject data (70.1% and 13.2%). Moreover, our models also show gains in cross-language calibration performance. This increase in accuracy suggests the feasibility of developing practical silent speech decoding systems, particularly for speech-impaired patients.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。