arXiv:2505.15355cs.CLcs.LG2025-05被引 3

用脑磁图解码说话时的音素对,发现主动发音比听或重放更易解码。

Decoding Phone Pairs from MEG Signals Across Speech Modalities

  • 对比17人数据,用多种模型解码15组音素对,分析不同语音模式下的脑信号。
  • 主动说话时音素解码准确率达76.6%,远高于听或重播时约51%。
  • 低频脑电波(δ和θ)贡献最大,适合研究语言生成与脑机接口应用。

理解语音产生背后的神经机制对于推进认知神经科学理论和开发实用通信技术至关重要。本研究利用脑磁图(MEG)信号,解码语音产生、被动聆听和语音回放任务中的音素信息。基于17名受试者的数据集,我们进行了15组音素对的成对分类。比较了正则化线性模型与神经网络等多种机器学习方法的性能。结果表明,主动说话时解码准确率高达76.6%,显著高于被动聆听和语音回放模式下的约51%,说明外显言语中包含更丰富的神经信息。在各类模型中,Elastic Net分类器持续优于复杂神经网络,凸显在高维、小样本的MEG数据上,传统正则化方法仍具优势。此外,特定脑电频率带分析显示,低频振荡(尤其是Delta: 0.2–3 Hz 和 Theta: 4–7 Hz)对解码准确性贡献最显著,提示这些频段编码了关键的语言生成神经过程。尽管采用先进去噪方法,仍难以排除残余肌肉或运动伪影的影响,表明需进一步改进方法。总体而言,研究强调了考察外显言语范式的重要性,其虽复杂,但为改善严重语言障碍者的脑机接口提供了可能。

原文摘要 · Abstract (English)

Understanding the neural mechanisms underlying speech production is essential for both advancing cognitive neuroscience theory and developing practical communication technologies. In this study, we investigated magnetoencephalography signals to decode phones from brain activity during speech production and perception (passive listening and voice playback) tasks. Using a dataset comprising 17 participants, we performed pairwise phone classification, extending our analysis to 15 phonetic pairs. Multiple machine learning approaches, including regularized linear models and neural network architectures, were compared to determine their effectiveness in decoding phonetic information. Our results demonstrate significantly higher decoding accuracy during speech production (76.6%) compared to passive listening and playback modalities (~51%), emphasizing the richer neural information available during overt speech. Among the models, the Elastic Net classifier consistently outperformed more complex neural networks, highlighting the effectiveness of traditional regularization techniques when applied to limited and high-dimensional MEG datasets. Besides, analysis of specific brain frequency bands revealed that low-frequency oscillations, particularly Delta (0.2-3 Hz) and Theta (4-7 Hz), contributed the most substantially to decoding accuracy, suggesting that these bands encode critical speech production-related neural processes. Despite using advanced denoising methods, it remains unclear whether decoding solely reflects neural activity or if residual muscular or movement artifacts also contributed, indicating the need for further methodological refinement. Overall, our findings underline the critical importance of examining overt speech production paradigms, which, despite their complexity, offer opportunities to improve brain-computer interfaces to help individuals with severe speech impairments.

脑机接口语音解码脑磁图音素识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。