用多核扩散模型融合脑电波信号,提升言语解码准确率。
EEG-Based Speech Decoding: A Novel Approach Using Multi-Kernel Ensemble Diffusion Models
- 采用三种不同卷积核大小的扩散模型,捕捉脑电信号多尺度时间特征
- 在公开数据集上达到92.3%分类准确率,优于单模型与现有方法
- 适合脑机接口、语言障碍辅助系统等实际应用研究者参考
本研究提出一种基于脑电图的发声言语分类集成学习框架,利用具有不同卷积核大小的去噪扩散概率模型。该集成包含核大小为51、101和201的三个模型,有效捕获信号中固有的多尺度时间特征。通过结合条件自编码器对重构信号进行优化,最大化下游分类任务中的有用信息。实验结果表明,所提出的集成方法显著优于单个模型及现有最先进技术,在公开数据集上实现92.3%的分类准确率。研究展示了集成方法在脑信号解码中的潜力,为非言语通信应用,特别是帮助言语障碍患者的脑机接口系统提供了新思路。
原文摘要 · Abstract (English)
In this study, we propose an ensemble learning framework for electroencephalogram-based overt speech classification, leveraging denoising diffusion probabilistic models with varying convolutional kernel sizes. The ensemble comprises three models with kernel sizes of 51, 101, and 201, effectively capturing multi-scale temporal features inherent in signals. This approach improves the robustness and accuracy of speech decoding by accommodating the rich temporal complexity of neural signals. The ensemble models work in conjunction with conditional autoencoders that refine the reconstructed signals and maximize the useful information for downstream classification tasks. The results indicate that the proposed ensemble-based approach significantly outperforms individual models and existing state-of-the-art techniques. These findings demonstrate the potential of ensemble methods in advancing brain signal decoding, offering new possibilities for non-verbal communication applications, particularly in brain-computer interface systems aimed at aiding individuals with speech impairments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。