arXiv:2501.08124eess.AScs.SD2025-01被引 4

用真实对话测试视听协同对语音追踪的提升效果

Neural Speech Tracking in a Virtual Acoustic Environment: Audio-Visual Benefit for Unscripted Continuous Speech

  • 在虚拟声学环境中用自然语流测试视听融合机制
  • 有口型时神经追踪准确率显著高于仅听声音,尤其在噪声中
  • 说话者音调、唇开度等个体差异影响视听整合效果

听觉-视觉协同效应在语音感知中已被广泛证实,尤其在听力困难或环境嘈杂时表现明显。然而,多数研究依赖受控实验室环境下预先设计的脚本化语音材料。本研究采用未经训练的说话人自然对话,在虚拟声学环境中,通过脑电图(EEG)和皮层语音追踪技术,比较了视听、仅听、仅看及遮蔽口型四种条件下的神经响应,以分离唇动的作用。同时分析了说话者个体特征(如音高、抖动、唇部开度)对视听整合的影响。结果显示,在背景噪声下,视听条件显著提升了语音追踪能力,而遮蔽口型的表现与仅听声音相当,表明唇动在不利听觉条件下至关重要。研究证实了使用自然语音进行皮层语音追踪的可行性,并揭示了说话者个体特征对真实场景中视听整合的关键影响。

原文摘要 · Abstract (English)

The audio visual benefit in speech perception, where congruent visual input enhances auditory processing, is well documented across age groups, particularly in challenging listening conditions and among individuals with varying hearing abilities. However, most studies rely on highly controlled laboratory environments with scripted stimuli. Here, we examine the audio visual benefit using unscripted, natural speech from untrained speakers within a virtual acoustic environment. Using electroencephalography (EEG) and cortical speech tracking, we assessed neural responses across audio visual, audio only, visual only, and masked lip conditions to isolate the role of lip movements. Additionally, we analysed individual differences in acoustic and visual features of the speakers, including pitch, jitter, and lip openness, to explore their influence on the audio visual speech tracking benefit. Results showed a significant audio visual enhancement in speech tracking with background noise, with the masked lip condition performing similarly to the audio-only condition, emphasizing the importance of lip movements in adverse listening situations. Our findings reveal the feasibility of cortical speech tracking with naturalistic stimuli and underscore the impact of individual speaker characteristics on audio-visual integration in real world listening contexts.

语音追踪视听融合自然语流

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。