用脑电活动重建高保真语音,为失语者提供新沟通可能。
NeuroIncept Decoder for High-Fidelity Speech Reconstruction from Neural Activity
- 融合高γ波段时频特征与混合CNN-GRU架构,解码神经信号
- 预测与实际声谱相关系数表现稳健,但个体差异明显
- 适合脑机接口、神经修复领域研究者参考
本文提出一种新型算法,用于从侵入式脑电图(EEG)记录的神经活动中合成语音。该系统为严重言语障碍患者提供了有前景的沟通解决方案。核心方法是将从脑电信号中计算出的高γ波段时频特征,与先进的NeuroIncept Decoder架构结合。该神经网络融合卷积神经网络(CNN)和门控循环单元(GRU),从神经模式重建音频声谱图。模型在预测与真实声谱间表现出稳健的均值相关系数,但跨被试差异表明参与者间存在不同的神经处理机制。研究展示了神经解码技术在恢复言语障碍者交流能力方面的潜力,并为脑机接口技术的未来发展铺平道路。
原文摘要 · Abstract (English)
This paper introduces a novel algorithm designed for speech synthesis from neural activity recordings obtained using invasive electroencephalography (EEG) techniques. The proposed system offers a promising communication solution for individuals with severe speech impairments. Central to our approach is the integration of time-frequency features in the high-gamma band computed from EEG recordings with an advanced NeuroIncept Decoder architecture. This neural network architecture combines Convolutional Neural Networks (CNNs) and Gated Recurrent Units (GRUs) to reconstruct audio spectrograms from neural patterns. Our model demonstrates robust mean correlation coefficients between predicted and actual spectrograms, though inter-subject variability indicates distinct neural processing mechanisms among participants. Overall, our study highlights the potential of neural decoding techniques to restore communicative abilities in individuals with speech disorders and paves the way for future advancements in brain-computer interface technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。