arXiv:2608.26213cs.SDcs.CV2026-08中稿 · Interspeech 2026

根据可靠性动态调整对比解码强度,提升语音识别在噪声下的鲁棒性。

Attention-Guided Reliability Scaling for Contrastive Decoding in Robust Audio-Visual Speech Recognition

论文配图:Attention-Guided Reliability Scaling for Contrastive Decoding in Robust Audio-Visual Speech Recognition
图 1 · 摘自论文原文
  • 基于注意力和模型间预测差异生成可靠性信号
  • 在LRS3数据集上清洁与低信噪比场景均提升性能
  • 适合需要抗噪声的多模态语音识别应用

基于大语言模型(LLM)的视听语音识别(AVSR)系统在噪声环境下表现稳健。对比解码(CD)通过推理时对比弱模型与强模型的输出来稳定生成,无需额外训练。本文将CD应用于AVSR,对比仅音频输入与全音视频输入在同一模型内的表现。然而,固定强度的对比会带来权衡:强干预在严重噪声下有益,但在干净条件下可能过度修正可靠预测。为此,我们提出可靠性感知的对比解码强度调节方法,不使用固定强度,而是根据注意力动态和跨模型预测差异生成的可靠性信号,自适应地调整每个词元的对比影响。在LRS3数据集上的实验表明,该方法在清洁及低信噪比条件下均实现一致提升。

原文摘要 · Abstract (English)

Large language model (LLM)-based audio-visual speech recognition (AVSR) systems are robust under noise. Contrastive decoding (CD), originally introduced to stabilize LLM generation by contrasting a weaker model against a stronger one at inference time, adjusts predictions without additional training. In this work, we apply CD to AVSR by contrasting audio-only conditioning with full audio-visual conditioning within the same underlying model. However, using a fixed contrastive strength introduces a trade-off across noise levels: stronger intervention helps under severe noise but may over-correct reliable predictions in clean conditions. We propose reliability-aware scaling of CD for AVSR. Instead of using a fixed strength, we adaptively modulate the contrastive influence at each token based on reliability signals derived from attention dynamics and inter-model predictive divergence. Experiments on LRS3 show consistent improvements across clean and low-SNR conditions.

视听识别对比解码可靠性感知鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。