测试大模型听觉能力,发现其常编造声音事件,提出逐步推理方法改进。
Can Large Audio-Language Models Truly Hear? Tackling Hallucinations with Multi-Task Assessment and Stepwise Audio Reasoning
- 设计三类任务评估模型对声音存在、顺序和属性的理解能力
- 实验显示模型在识别声音事件顺序上准确率不足60%
- 采用多轮思维链推理,显著提升模型在关键任务上的表现
近年来,大型音频-语言模型(LALMs)在理解与推理音频及语音信息方面展现出令人印象深刻的潜力。然而,这些模型仍面临幻觉问题,包括虚构不存在的声音事件、错误判断声音事件的顺序以及错误归因声音来源,严重损害其可靠性与实际应用价值。为系统评估这些问题,我们提出了三项不同任务:声音对象存在性、时间顺序与声音属性。这些任务旨在检验模型对关键音频信息的理解程度。实验结果揭示了模型在基础任务上的局限性,凸显出在识别特定声音事件、确定事件序列和识别声源方面仍需改进。为此,我们引入一种多轮思维链(multi-turn chain-of-thought)方法,该方法在所提任务中显著提升了模型性能。
原文摘要 · Abstract (English)
Recent advancements in large audio-language models (LALMs) have shown impressive capabilities in understanding and reasoning about audio and speech information. However, these models still face challenges, including hallucinating non-existent sound events, misidentifying the order of sound events, and incorrectly attributing sound sources, which undermine their reliability and real-world application. To systematically evaluate these issues, we propose three distinct tasks: object existence, temporal order, and object attribute within audio. These tasks assess the models' comprehension of critical audio information aspects. Our experimental results reveal limitations in these fundamental tasks, underscoring the need for better models in recognizing specific sound events, determining event sequences, and identifying sound sources. To improve performance in these areas, we introduce a multi-turn chain-of-thought approach, which demonstrates significantly improved model performance across the proposed tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。