arXiv:2506.02012cs.CVcs.SD2025-06被引 3

用大模型提升看口型识音效果,通过规模扩展、上下文引导和迭代优化实现突破

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing

  • 基于大模型规模实验发现口型识别存在可预测的性能增长规律
  • 引入上下文文本指导解码,使识别准确率显著提升
  • 通过多轮迭代修正输出,逐步降低识别错误率

视觉语音识别(VSR)通过分析嘴唇运动来转录语音。近期,大型语言模型(LLMs)被引入VSR系统,带来显著性能提升。然而,LLMs在该任务中的潜力尚未充分探索,如何有效利用仍不明确。本文系统研究了如何更好利用LLMs进行VSR,提出三项关键贡献:(1) 规模测试:研究LLM大小对VSR性能的影响,确认了该任务中存在缩放规律;(2) 上下文感知解码:在解码过程中加入上下文文本以引导生成,提高识别准确性;(3) 迭代精炼:提出迭代式优化生成结果的方法,逐步减少识别错误。大量实验证明,通过这些设计,可充分挖掘LLMs在VSR中的潜力,带来显著性能提升。

原文摘要 · Abstract (English)

Visual Speech Recognition (VSR) transcribes speech by analyzing lip movements. Recently, Large Language Models (LLMs) have been integrated into VSR systems, leading to notable performance improvements. However, the potential of LLMs has not been extensively studied, and how to effectively utilize LLMs in VSR tasks remains unexplored. This paper systematically explores how to better leverage LLMs for VSR tasks and provides three key contributions: (1) Scaling Test: We study how the LLM size affects VSR performance, confirming a scaling law in the VSR task. (2) Context-Aware Decoding: We add contextual text to guide the LLM decoding, improving recognition accuracy. (3) Iterative Polishing: We propose iteratively refining LLM outputs, progressively reducing recognition errors. Extensive experiments demonstrate that by these designs, the great potential of LLMs can be largely harnessed, leading to significant VSR performance improvement.

视觉语音识别大模型上下文解码迭代优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。