让唇语模型同时适应说话人视觉与语言习惯,提升真实场景识别率。
Personalized Lip Reading: Adapting to Your Unique Lip Movements with Vision and Language
- 双模态自适应:在视觉和语言层面同时调整预训练模型
- 新数据集支持10万词量、多姿态,实现真实场景句子级测试
- 比现有方法更有效,适合个性化唇语系统研发
唇语识别旨在通过分析唇部动作预测语音内容。尽管技术不断进步,但模型在面对未见过的说话人时性能下降,主要因其对唇形等视觉差异敏感。现有自适应方法仅关注视觉模态,未探索目标说话人词汇选择等语言信息的适配。此外,以往数据集词汇量小、姿态变化有限,难以验证方法在真实场景下的效果。为此,本文提出一种新型说话人自适应唇语识别方法,在视觉和语言双层面适配预训练模型,结合提示调优与LoRA技术实现高效微调。为验证其在真实场景中的有效性,我们构建了新数据集VoxLRS-SA,源自VoxCeleb2与LRS3,包含约10万个词汇、丰富姿态变化,首次实现英文场景下句子级唇语识别的野外验证。实验表明,现有自适应方法在野外场景中已能提升性能,而本文方法进一步取得更大改进。
原文摘要 · Abstract (English)
Lip reading aims to predict spoken language by analyzing lip movements. Despite advancements in lip reading technologies, performance degrades when models are applied to unseen speakers due to their sensitivity to variations in visual information such as lip appearances. To address this challenge, speaker adaptive lip reading technologies have advanced by focusing on effectively adapting a lip reading model to target speakers in the visual modality. However, the effectiveness of adapting language information, such as vocabulary choice, of the target speaker has not been explored in previous works. Additionally, existing datasets for speaker adaptation have limited vocabulary sizes and pose variations, which restrict the validation of previous speaker-adaptive methods in real-world scenarios. To address these issues, we propose a novel speaker-adaptive lip reading method that adapts a pre-trained model to target speakers at both vision and language levels. Specifically, we integrate prompt tuning and the LoRA approach, applying them to a pre-trained lip reading model to effectively adapt the model to target speakers. Furthermore, to validate its effectiveness in real-world scenarios, we introduce a new dataset, VoxLRS-SA, derived from VoxCeleb2 and LRS3. It contains a vocabulary of approximately 100K words, offers diverse pose variations, and enables the validation of adaptation methods in the wild, sentence-level lip reading for the first time in English. Through various experiments, we demonstrate that the existing speaker-adaptive method also improves performance in the wild at the sentence level. Moreover, we show that the proposed method achieves larger improvements compared to the previous works.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。