用双语音视觉模块提升音视频语音识别效率,降低模型参数。
DCIM-AVSR : Efficient Audio-Visual Speech Recognition via Dual Conformer Interaction Module
- 引入双协变交互模块,直接建模音视频层级关系。
- 参数量减少,预训练时选择性更新参数,提升效率。
- 适合资源受限场景下的实时音视频识别应用。
语音识别技术使机器能够理解并处理人类语言,将口语转换为文本或指令,广泛应用于虚拟助手、转录服务和通信工具中。音视频语音识别(AVSR)通过融合唇动、面部表情等视觉信息,在嘈杂环境中显著优于传统语音识别。尽管大规模数据训练的AVSR模型可达到甚至超过人类水平的准确率,但其高参数量带来了高昂的训练成本和部署难题。为此,本文提出一种高效AVSR模型,通过引入双协变交互模块(DCIM)有效减少参数数量;同时设计一种预训练方法,通过选择性更新参数进一步优化性能。与传统模型需独立学习音视频层级关系不同,本方法将该特性直接嵌入架构设计中,兼顾了效率与精度,为实际应用提供更可行的解决方案。
原文摘要 · Abstract (English)
Speech recognition is the technology that enables machines to interpret and process human speech, converting spoken language into text or commands. This technology is essential for applications such as virtual assistants, transcription services, and communication tools. The Audio-Visual Speech Recognition (AVSR) model enhances traditional speech recognition, particularly in noisy environments, by incorporating visual modalities like lip movements and facial expressions. While traditional AVSR models trained on large-scale datasets with numerous parameters can achieve remarkable accuracy, often surpassing human performance, they also come with high training costs and deployment challenges. To address these issues, we introduce an efficient AVSR model that reduces the number of parameters through the integration of a Dual Conformer Interaction Module (DCIM). In addition, we propose a pre-training method that further optimizes model performance by selectively updating parameters, leading to significant improvements in efficiency. Unlike conventional models that require the system to independently learn the hierarchical relationship between audio and visual modalities, our approach incorporates this distinction directly into the model architecture. This design enhances both efficiency and performance, resulting in a more practical and effective solution for AVSR tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。