测试大模型在复杂音频中识别事件的能力,发现越多事件越易出错。
A Sensitivity Analysis of Multi-Event Audio Grounding in Audio LLMs
- 构建两类查询:真实存在事件和不存在事件,评估模型准确性
- 多事件场景下准确率下降30%,误报率上升40%以上
- 适合关注音频大模型可靠性与幻觉问题的研究者
音频大模型虽具备强大听觉理解能力,但在复杂声景中的可靠性仍不明确。不同于以往小规模或受控查询的实验,本研究基于71,000个AudioCapsV2片段,提取标准化的(源,属性)事件,构建两类查询:真实事件查询用于真阳性检测,缺失事件查询用于探测幻觉生成,并在音频对齐文本嵌入空间中使用相似性过滤负样本。我们对四款SOTA音频大模型,采用12种提示变体,在每模型上进行50万次是/否查询评估。结果显示,随着事件数量增加,所有模型的真阳性率持续下降,假阳性率显著上升;提示设计引发显著的准确率与误报率权衡。置信度分析表明,多事件音频下模型不确定性增强,揭示改进空间。
原文摘要 · Abstract (English)
Audio LLMs have shown a strong ability to understand audio samples, yet their reliability in complex acoustic scenes remains under-explored. Unlike prior work limited to small scale or less controlled query construction, we present a large-scale evaluation of event grounding and false alarms as auditory scene complexity increases. Using 71K AudioCapsV2 clips, we extract normalized (source, attribute) events and build two query types: present-event queries for ground-truth detection and absent-event queries to probe hallucinations, using similarity-filtered negative sampling in an audio-aligned text embedding space. We evaluate four SOTA Audio LLMs with 12 prompt variants over 500K yes/no queries per model. Across models, increasing event count consistently lowers true-positive rate and raises false-positive rate, while prompts induce a strong trade-off between the two. Our confidence analysis shows that models become more uncertain on multi-event audio, revealing room for improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。