解析视觉发音单元在多模态语音识别中的编码机制
Interpreting the Role of Visemes in Audio-Visual Speech Recognition
- 用t-SNE可视化特征,发现视觉线索主导聚类
- 音频能细化模糊或罕见发音的特征表示
- 揭示多模态协同机制,适合模型优化研究者
多模态语音识别(AVSR)模型性能已超越纯音频模型,但其可解释性,尤其是视觉模态的作用,仍待深入。本文采用多种可解释性技术,分析当前领先的AVSR模型AV-HuBERT中发音单元(visemes)的编码方式。首先利用t-分布随机邻域嵌入(t-SNE)可视化特征,发现视觉线索驱动自然聚类,且音频进一步优化该结构。随后通过探针分析表明,音频对视觉模糊或低频出现的发音单元具有显著的特征修正作用。研究揭示了多模态间的动态交互机制,为提升视觉信息利用效率提供了新思路。
原文摘要 · Abstract (English)
Audio-Visual Speech Recognition (AVSR) models have surpassed their audio-only counterparts in terms of performance. However, the interpretability of AVSR systems, particularly the role of the visual modality, remains under-explored. In this paper, we apply several interpretability techniques to examine how visemes are encoded in AV-HuBERT a state-of-the-art AVSR model. First, we use t-distributed Stochastic Neighbour Embedding (t-SNE) to visualize learned features, revealing natural clustering driven by visual cues, which is further refined by the presence of audio. Then, we employ probing to show how audio contributes to refining feature representations, particularly for visemes that are visually ambiguous or under-represented. Our findings shed light on the interplay between modalities in AVSR and could point to new strategies for leveraging visual information to improve AVSR performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。