arXiv:2502.04834cs.CVcs.AI2025-02被引 1

轻量化操作让视频语音识别能在低功耗设备上运行

Lightweight Operations for Visual Speech Recognition

  • 设计高效运算模块,压缩模型体积同时保持识别能力
  • 在大规模数据集上实现低资源消耗与高准确率平衡
  • 适合移动端或边缘设备部署,代码模型开源可用

视觉语音识别(VSR)通过视频数据解码话语,在音频不可用时具有显著优势。然而,视频数据的高维特性导致计算开销巨大,需高性能硬件支持,限制了其在资源受限设备上的部署。本文通过设计高效的运算范式,构建轻量级但性能强大的VSR模型,显著降低资源需求并维持最小精度损失。我们在大规模公开数据集上训练和评估模型,验证其在实际应用中的有效性。通过大量消融实验,系统分析各模型的规模与复杂度。代码与训练好的模型将公开发布。

原文摘要 · Abstract (English)

Visual speech recognition (VSR), which decodes spoken words from video data, offers significant benefits, particularly when audio is unavailable. However, the high dimensionality of video data leads to prohibitive computational costs that demand powerful hardware, limiting VSR deployment on resource-constrained devices. This work addresses this limitation by developing lightweight VSR architectures. Leveraging efficient operation design paradigms, we create compact yet powerful models with reduced resource requirements and minimal accuracy loss. We train and evaluate our models on a large-scale public dataset for recognition of words from video sequences, demonstrating their effectiveness for practical applications. We also conduct an extensive array of ablative experiments to thoroughly analyze the size and complexity of each model. Code and trained models will be made publicly available.

视觉语音识别轻量化模型边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。