对比解码能有效纠正模型误判无音频的错误,但对自信误判无效。
How Contrastive Decoding Enhances Large Audio Language Models?
- 通过四种对比解码策略对比实验,发现音频感知与音频对比解码最有效。
- 对比解码可纠正因误判无音频或盲目猜测导致的错误,但无法修正错误推理。
- 提出转移矩阵框架,帮助判断哪些语音大模型适合用对比解码优化。
尽管对比解码(CD)已被证明能提升大型音频语言模型(LALMs)性能,但其内在机制及不同策略的有效性仍不明确。本研究系统评估了四种不同的CD策略在多种LALM架构上的表现。结果表明,音频感知解码和音频对比解码最为有效,但其效果因模型而异。为解释这种差异,我们引入转移矩阵框架,用于追踪推理过程中的错误模式变化。分析显示,CD能可靠修正模型错误地声称无音频或依赖不确定性猜测的错误;但对错误推理或自信误断则无效。最终,这些发现为根据基线错误特征选择最适合使用CD增强的LALM架构提供了清晰指导。
原文摘要 · Abstract (English)
While Contrastive Decoding (CD) has proven effective at enhancing Large Audio Language Models (LALMs), the underlying mechanisms driving its success and the comparative efficacy of different strategies remain unclear. This study systematically evaluates four distinct CD strategies across diverse LALM architectures. We identify Audio-Aware Decoding and Audio Contrastive Decoding as the most effective methods. However, their impact varies significantly by model. To explain this variability, we introduce a Transition Matrix framework to map error pattern shifts during inference. Our analysis demonstrates that CD reliably rectifies errors in which models falsely claim an absence of audio or resort to uncertainty-driven guessing. Conversely, it fails to correct flawed reasoning or confident misassertions. Ultimately, these findings provide a clear guideline for determining which LALM architectures are most suitable for CD enhancement based on their baseline error profiles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。