用语音大模型教视觉听觉模型,提升多模态语音识别效果
Audio-Visual Representation Learning via Knowledge Distillation from Speech Foundation Models
- 让语音大模型当老师,通过知识蒸馏指导视听模型学习
- 在多个任务上超越或持平现有最佳方法,性能提升显著
- 适合做语音识别、唇读等多模态研究的学者和工程师
音频-视觉表征学习对推进多模态语音处理任务(如唇读和视听语音识别)至关重要。近期,语音基础模型(SFMs)在各类语音相关任务中展现出出色的泛化能力。基于此,我们提出一种利用跨模态知识蒸馏的音频-视觉表征学习模型。该方法中,SFMs作为教师模型,通过干净音频输入提取多层隐藏表征;我们还引入多教师集成策略,对学生模型进行蒸馏训练,学生模型接收音视频数据输入。采用一种新型表征知识蒸馏损失函数,在预训练阶段及微调阶段均用于优化学生模型。实验使用了自监督的SFMs(WavLM)和有监督的SFMs(iFLYTEK-speech)。结果表明,所提方法在自动语音识别、视觉语音识别及视听语音识别任务中均达到优于或至少可比现有最先进基线的表现。此外,通过全面消融实验与学习表征可视化,验证了方法的有效性。
原文摘要 · Abstract (English)
Audio-visual representation learning is crucial for advancing multimodal speech processing tasks, such as lipreading and audio-visual speech recognition. Recently, speech foundation models (SFMs) have shown remarkable generalization capabilities across various speech-related tasks. Building on this progress, we propose an audio-visual representation learning model that leverages cross-modal knowledge distillation from SFMs. In our method, SFMs serve as teachers, from which multi-layer hidden representations are extracted using clean audio inputs. We also introduce a multi-teacher ensemble method to distill the student, which receives audio-visual data as inputs. A novel representational knowledge distillation loss is employed to train the student during pretraining, which is also applied during finetuning to further enhance the performance on downstream tasks. Our experiments utilized both a self-supervised SFM, WavLM, and a supervised SFM, iFLYTEK-speech. The results demonstrated that our proposed method achieved superior or at least comparable performance to previous state-of-the-art baselines across automatic speech recognition, visual speech recognition, and audio-visual speech recognition tasks. Additionally, comprehensive ablation studies and the visualization of learned representations were conducted to evaluate the effectiveness of our proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。