arXiv:2501.04359eess.AScs.CL2025-01被引 9

用VAE增强脑电数据,提升语音感知解码效果

Decoding EEG Speech Perception with Transformers and VAE-based Data Augmentation

  • 用变分自编码器生成合成脑电信号以扩充数据
  • 序列到序列模型在生成句子任务上优于分类模型
  • 适合脑机接口与言语障碍辅助技术研究者

从非侵入性脑信号(如脑电图,EEG)中解码语音,有望推动脑机接口发展,应用于无声通信及言语障碍者的辅助技术。然而,EEG语音解码面临数据噪声大、数据集有限、复杂任务性能差等挑战。本研究通过变分自编码器(VAEs)进行EEG数据增强以提升数据质量,并采用在肌电(EMG)任务中表现优异的先进序列到序列深度学习架构,应用于EEG语音解码。此外,还将其适配用于单词分类任务。基于Brennan数据集(包含受试者聆听叙述性语音时的EEG记录),我们预处理数据并评估分类与序列到序列模型在EEG转文字/句子任务中的表现。实验表明,VAEs具备重建人工EEG数据用于数据增强的潜力;序列到序列模型在句子生成任务中表现更优,尽管两类任务仍具挑战性。这些发现为未来EEG语音感知解码研究奠定基础,未来可拓展至无声或想象语音等语音生成任务。

原文摘要 · Abstract (English)

Decoding speech from non-invasive brain signals, such as electroencephalography (EEG), has the potential to advance brain-computer interfaces (BCIs), with applications in silent communication and assistive technologies for individuals with speech impairments. However, EEG-based speech decoding faces major challenges, such as noisy data, limited datasets, and poor performance on complex tasks like speech perception. This study attempts to address these challenges by employing variational autoencoders (VAEs) for EEG data augmentation to improve data quality and applying a state-of-the-art (SOTA) sequence-to-sequence deep learning architecture, originally successful in electromyography (EMG) tasks, to EEG-based speech decoding. Additionally, we adapt this architecture for word classification tasks. Using the Brennan dataset, which contains EEG recordings of subjects listening to narrated speech, we preprocess the data and evaluate both classification and sequence-to-sequence models for EEG-to-words/sentences tasks. Our experiments show that VAEs have the potential to reconstruct artificial EEG data for augmentation. Meanwhile, our sequence-to-sequence model achieves more promising performance in generating sentences compared to our classification model, though both remain challenging tasks. These findings lay the groundwork for future research on EEG speech perception decoding, with possible extensions to speech production tasks such as silent or imagined speech.

脑机接口语音解码VAEEEG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。