用噪声提示引导音频大模型减少幻觉,无需微调即可提升可靠性。
Noise-Aware In-Context Learning for Hallucination Mitigation in ALLMs
- 构建噪声先验库,通过上下文引入相关噪声样本
- 在音频字幕任务中将幻觉率从26.53%降至16.98%
- 适合关注音频生成可靠性的研究者和开发者
听觉大语言模型(ALLMs)在音频理解与推理任务中展现出强大泛化能力,但其可靠性仍受幻觉问题困扰。现有幻觉评估方法多为二分类任务,难以刻画生成任务中复杂的幻觉模式。当前缓解策略依赖微调,计算成本高。为此,我们提出即插即用的噪声感知上下文学习(NAICL)方法:构建噪声先验库,检索与输入音频相关的噪声示例并作为上下文先验,引导模型在声学证据不足时减少推测性关联,采取更保守的生成策略。此外,我们建立了音频字幕任务的幻觉基准,包括构建Clotho-1K多事件数据集,定义四类听觉幻觉,并引入幻觉类型分布等指标支持细粒度分析。实验表明,所有评测的ALLMs均表现出相似幻觉行为。所提NAICL方法将整体幻觉率从26.53%降低至16.98%。
原文摘要 · Abstract (English)
Auditory large language models (ALLMs) have demonstrated strong general capabilities in audio understanding and reasoning tasks. However, their reliability is still undermined by hallucination issues. Existing hallucination evaluation methods are formulated as binary classification tasks, which are insufficient to characterize the more complex hallucination patterns that arise in generative tasks. Moreover, current hallucination mitigation strategies rely on fine-tuning, resulting in high computational costs. To address the above limitations, we propose a plug-and-play Noise-Aware In-Context Learning (NAICL) method. Specifically, we construct a noise prior library, retrieve noise examples relevant to the input audio, and incorporate them as contextual priors, thereby guiding the model to reduce speculative associations when acoustic evidence is insufficient and to adopt a more conservative generation strategy. In addition, we establish a hallucination benchmark for audio caption tasks including the construction of the Clotho-1K multi-event benchmark dataset, the definition of four types of auditory hallucinations, and the introduction of metrics such as hallucination type distribution to support fine-grained analysis. Experimental results show that all evaluated ALLMs exhibit same hallucination behaviors. Moreover, the proposed NAICL method reduces the overall hallucination rate from 26.53% to 16.98%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。