arXiv:2411.11308eess.ASeess.SP2024-11被引 1

解析语音理解中声学与语义线索的作用,揭示注意力下的大脑处理机制。

Uncovering the role of semantic and acoustic cues in normal and dichotic listening

  • 设计匹配-不匹配任务,用深度学习模型融合声学、语义与脑电数据
  • 自然聆听时语义与声学线索表现相当,分耳听时语义优势显著
  • 首次量化证明右耳优势,适合认知神经科学与听觉研究者

语音理解是健康人脑的无意识任务,但其内在机制仍不清晰。本文旨在量化复杂听觉条件下声学与语义信息流的作用。提出一种基于脑电图(EEG)数据的匹配-不匹配(MM)分类范式,以分析语音线索的编码。构建多模态深度学习序列模型STEM,输入包括语音包络(声学)、文本表示(语义)和神经响应(EEG)。在两种条件上进行实验:一为自然被动聆听,二为需注意的分耳听任务。以MM任务为分析框架,发现:(a)语音感知按词边界碎片化;(b)自然聆听中声学与语义线索表现相似;(c)分耳听任务中语义线索显著优于声学线索。与先前模型相比,STEM性能显著提升。本研究量化了声学与语义线索在不同听觉任务中的作用,进一步支持分耳听中的右耳优势现象。

原文摘要 · Abstract (English)

Speech comprehension is an involuntary task for the healthy human brain, yet the understanding of the mechanisms underlying this brain functionality remains obscure. In this paper, we aim to quantify the role of acoustic and semantic information streams in complex listening conditions. We propose a paradigm to understand the encoding of the speech cues in electroencephalogram (EEG) data, by designing a match-mismatch (MM) classification task. The MM task involves identifying whether the stimulus (speech) and response (EEG) correspond to each other. We build a multimodal deep-learning based sequence model STEM, which is input with acoustic stimulus (speech envelope), semantic stimulus (textual representations of speech), and the neural response (EEG data). We perform extensive experiments on two separate conditions, i) natural passive listening and, ii) a dichotic listening requiring auditory attention. Using the MM task as the analysis framework, we observe that - a) speech perception is fragmented based on word boundaries, b) acoustic and semantic cues offer similar levels of MM task performance in natural listening conditions, and c) semantic cues offer significantly improved MM classification over acoustic cues in dichotic listening task. The comparison of the STEM with previously proposed MM models shows significant performance improvements for the proposed approach. The analysis and understanding from this study allows the quantification of the roles played by acoustic and semantic cues in diverse listening tasks and in providing further evidences of right-ear advantage in dichotic listening.

听觉认知脑电分析多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。