arXiv:2506.20361eess.AScs.SD2025-06中稿 · Interspeech 2025

对比音视频模型与纯音频模型,发现前者未捕捉到语音感知的时序差异。

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models

  • 用线性分类器追踪音视频和纯音频模型的发音信息解码时间。
  • 音视频模型在语音信息可用上比纯音频模型提前约20毫秒。
  • 因时间分辨率低,音视频模型未能真实反映多模态语音感知的动态过程。

人类语音感知具有多模态特性。自然语音中,唇部动作可能先于相应发声出现100-300毫秒,尤其对特定辅音,这会影响人类听者神经发音编码的时间进程。然而,自监督学习模型是否能捕捉音频与视觉线索之间的这种非可忽略的时序差异仍不清楚。我们比较了音视频版AV-HuBERT与纯音频版HuBERT,使用线性分类器追踪其嵌入表示中的发音信息解码时间。结果显示,音视频模型的音素信息仅在约20毫秒前即已可被解码,可能源于其较低的时间分辨率及特征拼接机制。这表明AV-HuBERT未能充分捕捉多模态语音感知的时间动态,限制了其在建模多模态语音感知过程中的适用性。

原文摘要 · Abstract (English)

Human speech perception is multimodal. In natural speech, lip movements can precede corresponding voicing by a non-negligible gap of 100-300 ms, especially for specific consonants, affecting the time course of neural phonetic encoding in human listeners. However, it remains unexplored whether self-supervised learning models, which have been used to simulate audio-visual integration in humans, can capture this asynchronicity between audio and visual cues. We compared AV-HuBERT, an audio-visual model, with audio-only HuBERT, by using linear classifiers to track their phonetic decodability over time. We found that phoneme information becomes available in AV-HuBERT embeddings only about 20 ms before HuBERT, likely due to AV-HuBERT's lower temporal resolution and feature concatenation process. It suggests AV-HuBERT does not adequately capture the temporal dynamics of multimodal speech perception, limiting its suitability for modeling the multimodal speech perception process.

自监督学习音视频融合语音感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。