arXiv:2501.06326cs.LGeess.IV2025-01被引 3

用脑电图解码说话内容,少标签也能高精度。

On Creating A Brain-To-Text Decoder

  • 直接用原始脑电信号解码语言,无需复杂预处理。
  • 在有限标注数据下,词错误率低于现有方法。
  • 揭示电极密度和词汇量对解码效果的关键影响。

脑解码已成为神经科学中快速发展的核心技术。本文聚焦于使用原始脑电图(EEG)信号解码大脑活动,提供更高效的方法以深化对人脑的理解。研究重点评估脑机接口(BCI)在语音产生神经信号解码中的表现,特别关注词汇量、电极密度和训练数据对系统性能的影响。实验表明,在Librispeech基准上,通过语音处理的无标签数据预训练可实现具有竞争力的词错误率(WER)。此外,在标注数据极少的情况下,该方法仍显著优于先前最优技术,且所需标签数量大幅减少。研究还系统分析了语音识别中的错误模式,以及模型规模和无标签训练数据的影响。结果强调了词汇量和电极密度对提升BCI性能的重要性,建议增加微电极数量并优化语言模型。

原文摘要 · Abstract (English)

Brain decoding has emerged as a rapidly advancing and extensively utilized technique within neuroscience. This paper centers on the application of raw electroencephalogram (EEG) signals for decoding human brain activity, offering a more expedited and efficient methodology for enhancing our understanding of the human brain. The investigation specifically scrutinizes the efficacy of brain-computer interfaces (BCI) in deciphering neural signals associated with speech production, with particular emphasis on the impact of vocabulary size, electrode density, and training data on the framework's performance. The study reveals the competitive word error rates (WERs) achievable on the Librispeech benchmark through pre-training on unlabelled data for speech processing. Furthermore, the study evaluates the efficacy of voice recognition under configurations with limited labeled data, surpassing previous state-of-the-art techniques while utilizing significantly fewer labels. Additionally, the research provides a comprehensive analysis of error patterns in voice recognition and the influence of model size and unlabelled training data. It underscores the significance of factors such as vocabulary size and electrode density in enhancing BCI performance, advocating for an increase in microelectrodes and refinement of language models.

脑机接口语音解码少样本学习脑电图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。