arXiv:2505.19626cs.SDeess.AS2025-05

用脑电图解码汉语发音的相对音高,揭示大脑如何忽略说话人差异。

Decoding Speaker-Normalized Pitch from EEG for Mandarin Perception

  • 基于脑电数据直接解码音高轮廓,采用CE-ViViT模型提升精度。
  • 标准化音高解码误差更小,证明大脑编码的是相对音高而非绝对频率。
  • 适用于语音感知、脑机接口与跨语言研究者参考。

同一段语音内容由不同说话人发出时,音高轮廓差异显著,但听者语义理解不受影响。这可能源于大脑对音高的感知独立于个体的音高范围。本文通过记录受试者聆听带不同声调、音素和说话人的普通话单字时的脑电图(EEG),提出CE-ViViT模型,直接从脑电数据中解码原始或说话人标准化的音高轮廓。实验表明,该模型能以较小误差实现音高轮廓解码,性能达到当前最先进的脑电回归方法水平。更重要的是,说话人标准化音高轮廓的解码更为准确,支持神经编码中存在相对音高的机制。

原文摘要 · Abstract (English)

The same speech content produced by different speakers exhibits significant differences in pitch contour, yet listeners' semantic perception remains unaffected. This phenomenon may stem from the brain's perception of pitch contours being independent of individual speakers' pitch ranges. In this work, we recorded electroencephalogram (EEG) while participants listened to Mandarin monosyllables with varying tones, phonemes, and speakers. The CE-ViViT model is proposed to decode raw or speaker-normalized pitch contours directly from EEG. Experimental results demonstrate that the proposed model can decode pitch contours with modest errors, achieving performance comparable to state-of-the-art EEG regression methods. Moreover, speaker-normalized pitch contours were decoded more accurately, supporting the neural encoding of relative pitch.

脑电解码语音感知相对音高中文语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。