arXiv:2510.12851cs.SDcs.LG2025-10中稿 · Interspeech 2026被引 5

用静音对比抑制音频大模型幻觉,提升问答准确率。

Silence is Golden: Mitigating Hallucinations in Large Audio-Language Models via Layer-Weighted Vector Steering

  • 以静音为基准,通过对比抑制幻觉输出。
  • 在Gemma上召回率提升15.6%,达69.0%。
  • 不需训练,适合想增强音频理解的开发者。

大型音频-语言模型(LALMs)在音频问答任务中表现优异,但常产生脱离音频内容的幻觉。据我们所知,这是首个将向量引导技术应用于音频领域的方法。不同于文本引导,本文提出基于静音锚定的对比策略,通过将活跃音频与静音基线对比,引导模型远离幻觉。探针分析显示特定层表征与输出正确性高度相关。据此提出无需训练的分层加权向量引导(LWVS),在关键层增强引导强度。在音频幻觉问答数据集上,LWVS显著优于基线,使Gemma模型召回率从53.4%提升至69.0%(+15.6%)。重要的是,MMAU基准测试表明,LWVS不仅保留且提升了通用音频理解能力,在Qwen模型上实现相对8%的准确率提升(54.8%→59.2%)。

原文摘要 · Abstract (English)

Large Audio-Language Models (LALMs) excel in Audio QA but often suffer from hallucinations ungrounded in the audio. To our knowledge, we are the first to propose applying vector steering to the audio domain to mitigate this. Unlike text-based steering, our silence-anchored contrastive approach steers the model away from hallucinations by contrasting active audio against a silent baseline. Probing internal states reveals a strong correlation between specific layer representations and output correctness. Leveraging this, we introduce Layer-Weighted Vector Steering (LWVS), a training-free intervention that increases steering strength at influential layers. On the Audio Hallucination QA dataset, LWVS significantly outperforms baselines, boosting Recall on the Gemma model by 15.6% (53.4% to 69.0%). Crucially, MMAU benchmark tests confirm LWVS preserves and even enhances general audio understanding, achieving an 8% relative accuracy increase on the Qwen model (54.8% to 59.2%).

音频生成幻觉抑制向量引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。