arXiv:2510.22961eess.AS2025-10被引 1

用大模型统一处理语音、视觉和音视频识别,让预训练语音模型轻松适配多模态场景。

Adapting Speech Foundation Models for Unified Multimodal Speech Recognition with Large Language Models

  • 将视觉信息注入冻结的语音模型,通过大模型实现多模态融合与统一表征。
  • 在清洁和嘈杂环境下均超越现有最佳方法,三类任务性能全面领先。
  • 训练策略通用性强,可适配不同语音模型和大模型,适合多模态语音研究者。

尽管语音基础模型(SFMs)在纯音频任务中表现卓越,但其在多模态场景下的适应性仍待探索。本文提出UASR-LLM框架,利用大语言模型(LLMs)作为文本解码器,将冻结的SFMs统一适配于视觉语音识别(VSR)、自动语音识别(ASR)和音视频语音识别(AVSR)。通过视觉注入模块将视觉表征注入多个SFM层,实现多模态融合与统一表征学习。增强后的SFMs通过前馈适配器连接至仅解码器的LLMs,由拼接表征与指令提示引导转录。提出两阶段训练策略:第一阶段为视觉注入预训练,对齐音频、视觉及音视频表征;第二阶段为语音识别微调,整合LLMs实现跨任务统一优化。实验表明,在清洁与噪声条件下,该方法在三类任务上均显著优于现有最先进基线。消融实验进一步验证了该训练策略在多种SFMs与LLMs间的泛化能力。

原文摘要 · Abstract (English)

While speech foundation models (SFMs) have demonstrated remarkable performance in audio-only tasks, their adaptation to multimodal scenarios remains underexplored. This work presents UASR-LLM, a novel framework that adapts frozen SFMs to unified visual speech recognition (VSR), automatic speech recognition (ASR), and audio-visual speech recognition (AVSR) by leveraging large language models (LLMs) as text decoders. Visual representations are injected into multiple SFM layers via visual injection modules, enabling multimodal fusion and unified representation learning. The augmented SFMs are connected to decoder-only LLMs through a feed-forward adaptor, where concatenated representations and instruction prompts guide transcription. We propose a two-stage training strategy consisting of visual injection pretraining followed by speech recognition finetuning. The pretraining stage aligns audio, visual, and audio-visual representations within the frozen SFM backbone, while the finetuning stage integrates LLMs for unified optimization across speech recognition tasks. Experimental results demonstrate superior performance over state-of-the-art baselines across VSR, ASR, and AVSR under both clean and noisy conditions. Ablation studies further confirm generalization across various SFMs and LLMs, validating the effectiveness of the proposed training strategy.

多模态语音识别大模型视觉语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。