轻量级视觉语音识别模型,准确率显著提升。
Enhancing CTC-Based Visual Speech Recognition
- 基于ASR模型知识蒸馏,优化视频预处理与特征归一化
- 在LRS2/LRS3上性能超越现有CTC模型,不增算力与数据
- 适合资源受限场景的高效语音识别应用
本文提出LiteVSR2,是此前高效视觉语音识别方法的改进版本。基于预训练自动语音识别(ASR)模型的知识蒸馏框架,引入两项关键改进:稳定的视频预处理技术与蒸馏过程中的特征归一化。这两项改进在LRS2和LRS3基准上带来显著性能提升,使LiteVSR2成为当前无需增加训练数据或计算资源的最优CTC-based VSR模型。此外,我们通过不同模型复杂度和训练数据量的实验,验证了该方法的可扩展性。LiteVSR2在保持前代效率的同时大幅提升准确率,展示了资源高效提升视觉语音识别技术的潜力。
原文摘要 · Abstract (English)
This paper presents LiteVSR2, an enhanced version of our previously introduced efficient approach to Visual Speech Recognition (VSR). Building upon our knowledge distillation framework from a pre-trained Automatic Speech Recognition (ASR) model, we introduce two key improvements: a stabilized video preprocessing technique and feature normalization in the distillation process. These improvements yield substantial performance gains on the LRS2 and LRS3 benchmarks, positioning LiteVSR2 as the current best CTC-based VSR model without increasing the volume of training data or computational resources utilized. Furthermore, we explore the scalability of our approach by examining performance metrics across varying model complexities and training data volumes. LiteVSR2 maintains the efficiency of its predecessor while significantly enhancing accuracy, thereby demonstrating the potential for resource-efficient advancements in VSR technology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。