arXiv:2509.14880cs.SD2025-09中稿 · publication in ICA…被引 3

用大模型提升语音识别,发现关键在语言推理而非视觉理解

From Hype to Insight: Rethinking Large Language Model Integration in Visual Speech Recognition

  • 冻结或部分更新视觉编码器,测试不同大模型解码器表现
  • 合并多数据集训练后,错误率降至24.7%(LRS3)和47.0%(WildVSR)
  • 适合关注视觉语音识别中语言模型作用的研究者

自监督编码器的进步提升了视觉语音识别(VSR)性能。近期将这些编码器与大语言模型(LLM)解码器结合的方法虽提高了转录准确率,但尚不清楚其提升源于视觉理解还是更强的语言建模能力。本文系统评估了LLM解码器,通过冻结或选择性更新视觉编码器、扩展解码器规模、比较适配策略与架构,并在LRS2、LRS3及其组合数据集上进行训练。在LRS2、LRS3和WildVSR上的评估表明,扩大模型和优化适配带来的改进有限,而数据集合并显著增强泛化能力。语义分析显示,性能提升主要来自词汇层面而非语义处理。我们基于合并数据集训练的Llama-2-13B模型在LRS3上达到24.7%的词错误率(WER),在WildVSR上为47.0%,成为无额外监督条件下最先进模型。研究结果表明,LLM解码器主要提升上下文推理能力,而非改善视觉特征,强调需更强的视觉编码器以实现实质性进展。

原文摘要 · Abstract (English)

Advances in self-supervised encoders have improved Visual Speech Recognition (VSR). Recent approaches integrating these encoders with LLM decoders improves transcription accuracy; however, it remains unclear whether these gains stem from visual understanding or stronger language modeling. In this work, we systematically evaluate LLM decoders by freezing or selectively updating the visual encoder, scaling decoder size, comparing adaptation strategies and architectures, and varying training data across LRS2, LRS3, and their combination. Evaluation on LRS2, LRS3, and WildVSR shows that scaling and adaptation yield limited improvements, while combining datasets enhances generalization. Semantic analysis reveals that gains arise primarily from lexical rather than semantic processing. Our Llama-2-13B model trained on the combined set achieves 24.7% WER on LRS3 and 47.0% on WildVSR, establishing SOTA among models trained without additional supervision. Our findings indicate LLM decoders refine contextual reasoning rather than visual features, emphasizing the need for stronger visual encoders to drive meaningful progress.

视觉语音识别大语言模型自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。