用脑电波解码听、说、想三种言语状态,非侵入式更实用
Towards Unified Neural Decoding of Perceived, Spoken and Imagined Speech from EEG Signals
- 用深度学习区分听、说、默念、想象四种言语模式
- 伽马频段表现最佳,θ频段想象言语差异显著
- 为瘫痪患者提供新接口,适合神经工程与临床应用
大脑信号包含人类行为与心智图像的多种信息,对理解人类意图至关重要。脑机接口技术利用这些脑活动生成外部指令以控制环境,为瘫痪或闭锁综合征患者带来重要帮助。在脑机接口领域,脑到语音研究备受关注,聚焦于从脑信号直接合成可听语音。现有研究多依赖侵入式方法,且集中于口语数据。然而,人类表达存在多种言语状态,通过非侵入方式区分这些状态仍具挑战。本研究探究深度学习模型在非侵入式神经信号解码中的有效性,重点区分感知、外显、耳语及想象言语,涵盖多个频率带。基于空间卷积神经网络模块的模型在伽马频段表现最优;此外,θ频段的想象言语表现出与其它言语模式的统计显著差异。
原文摘要 · Abstract (English)
Brain signals accompany various information relevant to human actions and mental imagery, making them crucial to interpreting and understanding human intentions. Brain-computer interface technology leverages this brain activity to generate external commands for controlling the environment, offering critical advantages to individuals with paralysis or locked-in syndrome. Within the brain-computer interface domain, brain-to-speech research has gained attention, focusing on the direct synthesis of audible speech from brain signals. Most current studies decode speech from brain activity using invasive techniques and emphasize spoken speech data. However, humans express various speech states, and distinguishing these states through non-invasive approaches remains a significant yet challenging task. This research investigated the effectiveness of deep learning models for non-invasive-based neural signal decoding, with an emphasis on distinguishing between different speech paradigms, including perceived, overt, whispered, and imagined speech, across multiple frequency bands. The model utilizing the spatial conventional neural network module demonstrated superior performance compared to other models, especially in the gamma band. Additionally, imagined speech in the theta frequency band, where deep learning also showed strong effects, exhibited statistically significant differences compared to the other speech paradigms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。