arXiv:2507.18863cs.CVcs.CL2025-07中稿 · ICASSP 2026被引 3

通过音素级融合与大模型重建,提升视觉语音识别准确率

Phoneme-Level Visual Speech Recognition via Point-Visual Fusion and Language Model Reconstruction

  • 分两阶段:先预测音素,再用大模型重构为单词
  • 在LRS2和LRS3上分别达到17.4%和21.0%的词错误率
  • 有效缓解音素视觉模糊问题,适合低数据场景应用

视觉自动语音识别(V-ASR)是一项挑战性任务,仅依靠唇部运动和面部表情等视觉信息来理解口语。由于缺乏听觉线索以及音素间存在视觉相似性(即同形音),该任务尤为困难。现有方法通常直接从视觉信号预测词或字符,但常因同形音模糊导致高错误率,且需大量预训练数据。本文提出一种基于音素的两阶段新框架:第一阶段融合视觉特征与面部关键点运动,输出预测音素,降低训练复杂度;第二阶段使用NLLB编码器-解码器语言模型,将音素重构为完整词汇。结合大规模视觉数据微调,所提方法在LRS2和LRS3数据集上分别取得17.4%和21.0%的词错误率(WER),显著优于现有方法。

原文摘要 · Abstract (English)

Visual Automatic Speech Recognition (V-ASR) is a challenging task that involves interpreting spoken language solely from visual information, such as lip movements and facial expressions. This task is notably challenging due to the absence of auditory cues and the visual ambiguity of phonemes that exhibit similar visemes-distinct sounds that appear identical in lip motions. Existing methods often aim to predict words or characters directly from visual cues, but they commonly suffer from high error rates due to viseme ambiguity and require large amounts of pre-training data. We propose a novel phoneme-based two-stage framework that fuses visual and landmark motion features, followed by an LLM model for word reconstruction to address these challenges. Stage 1 consists of V-ASR, which outputs the predicted phonemes, thereby reducing training complexity. Meanwhile, the facial landmark features address speaker-specific facial characteristics. Stage 2 comprises an encoder-decoder LLM model, NLLB, that reconstructs the output phonemes back to words. Besides using a large visual dataset for deep learning fine-tuning, our PV-ASR method demonstrates superior performance by achieving 17.4% WER on the LRS2 and 21.0% WER on the LRS3 dataset.

视觉语音识别音素级建模大模型重建同形音模糊

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。