arXiv:2607.09001cs.SDeess.AS2026-07

用最优传输对齐多模态表示,提升语音识别精度。

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition

论文配图:Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition
图 1 · 摘自论文原文
  • 通过最优传输将音视频特征对齐至语言模型语义空间
  • 在多种信噪比下均超越现有方法,达最新性能
  • 适合关注多模态融合与语音识别的学者

基于大语言模型(LLM)的音视频语音识别(LLM-AVSR)近年来在恶劣声学环境下展现出强鲁棒性,通过融合互补的音视频信息实现。现有方法通常使用独立预训练的音频和视觉编码器,其输出被投影并融合为软提示,以条件化LLM进行语音识别。然而,大多数方法在多模态融合中未显式解决音频、视觉与文本模态间的表征差异,可能限制跨模态整合效果。本文提出一种基于最优传输(OT)的语义对齐框架用于LLM-AVSR。该方法在多模态融合前,显式地将音频与视觉表示对齐至LLM的语言嵌入空间。具体而言,利用OT估计概率耦合矩阵,刻画模态特异性特征与语言嵌入之间的结构化对应关系。所得OT耦合进一步作为软伪标签,监督对比学习,促使提取语义一致且跨模态一致的音视频表示。通过将多模态特征锚定于LLM的语言空间,所提框架促进更有效的多模态融合与解码。实验采用基于Whisper的音频编码器、基于AV-HuBERT的视觉编码器及LLaMA3.2-3B解码器,在LRS3-TED基准上验证,无论在干净还是噪声环境下,均持续优于强基线,且在广泛信噪比(SNRs)范围内达到当前最优性能。

原文摘要 · Abstract (English)

Large language model (LLM)-based audio-visual speech recognition (LLM-AVSR) has recently demonstrated strong robustness in adverse acoustic environments by leveraging complementary audio and visual information. Existing approaches typically employ independently pretrained acoustic and visual encoders, whose outputs are projected and fused as soft prompts to condition an LLM for speech recognition. However, most methods perform multimodal fusion without explicitly addressing the representational discrepancy between audio, visual and text modalities, potentially limiting the effectiveness of cross-modal integration. In this paper, we propose an optimal transport (OT)-based semantic alignment framework for LLM-AVSR. The proposed method explicitly bridges the modality gap by aligning the acoustic and visual representations with reference to the linguistic embedding space of the LLM before multimodal fusion. Specifically, OT is used to estimate probabilistic coupling matrices that characterize structured correspondences between modality-specific features and linguistic embeddings. The resulting OT couplings are further utilized as soft pseudo-labels to supervise contrastive learning, encouraging the extraction of semantically coherent and cross-modal consistent audio-visual representations. By anchoring multimodal features to the linguistic space of the LLM, the proposed framework facilitates more effective multimodal fusion and decoding. We implement the proposed framework using a Whisper-based acoustic encoder, an AV-HuBERT-based visual encoder, and a LLaMA3.2-3B decoder. Experiments conducted on the LRS3-TED benchmark demonstrate consistent improvements over strong baselines and achieve state-of-the-art performance under both clean and noisy evaluation conditions across a wide range of signal-to-noise ratios (SNRs).

语音识别多模态最优传输LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。