arXiv:2502.05758eess.AS2025-02中稿 · ESWA journal被引 7

跨语言迁移+目标说话人适配,提升中文唇读准确率

Target Speaker Lipreading by Audio-Visual Self-Distillation Pretraining and Speaker Adaptation

  • 用多模态自蒸馏预训练模型跨语言迁移
  • 针对特定说话人优化,降低字符错误率至77.3%
  • 融合面部与唇区输入,提升模型泛化能力

唇读技术在嘈杂环境下对人机交互具有重要意义。此前我们提出的自监督方法AV2vec利用多模态自蒸馏,在英语LRS3数据集上实现了出色的无说话人依赖唇读性能。但该方法存在训练成本高、非英语语言(如中文)音视频数据稀缺等问题,且多数研究集中于无说话人依赖模型,难以应对不同说话人之间的显著差异。为此,我们提出综合解决方案:首先,探索跨语言迁移学习,将源语言预训练的AV2vec模型适配到目标语言的唇读任务;其次,引入说话人适配策略,提升对特定目标说话人的识别精度,该方向此前研究较少;第三,通过分析唇区感兴趣区域(ROI)与全脸输入的互补性能,提出模型融合策略,显著提升整体表现。本方法在ChatCLR数据集测试集上达到77.3%的字符错误率(CER),优于2024年聊天场景中文唇读挑战赛的最优结果。

原文摘要 · Abstract (English)

Lipreading is an important technique for facilitating human-computer interaction in noisy environments. Our previously developed self-supervised learning method, AV2vec, which leverages multimodal self-distillation, has demonstrated promising performance in speaker-independent lipreading on the English LRS3 dataset. However, AV2vec faces challenges such as high training costs and a potential scarcity of audio-visual data for lipreading in languages other than English, such as Chinese. Additionally, most studies concentrate on speakerindependent lipreading models, which struggle to account for the substantial variation in speaking styles across di?erent speakers. To address these issues, we propose a comprehensive approach. First, we investigate cross-lingual transfer learning, adapting a pre-trained AV2vec model from a source language and optimizing it for the lipreading task in a target language. Second, we enhance the accuracy of lipreading for specific target speakers through a speaker adaptation strategy, which is not extensively explored in previous research. Third, after analyzing the complementary performance of lipreading with lip region-of-interest (ROI) and face inputs, we introduce a model ensembling strategy that integrates both, signi?cantly boosting model performance. Our method achieved a character error rate (CER) of 77.3% on the evaluation set of the ChatCLR dataset, which is lower than the top result from the 2024 Chat-scenario Chinese Lipreading Challenge.

唇读跨语言说话人适配多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。