arXiv:2505.13115cs.CLcs.AI2025-05中稿 · INTERSPEECH, 2025,…被引 11

构建音频时间推理数据集,评估大音频模型表现并提出不确定性度量。

Benchmarking and Confidence Evaluation of LALMs For Temporal Reasoning

  • 设计TREA数据集,专用于测试音频语言模型的时间推理能力。
  • 开源模型在任务上普遍低于人类表现,差距显著。
  • 引入输入语义扰动不变性度量,揭示准确率与不确定性的分离现象。

大型语言模型(LLM)的成功推动了多模态领域将视觉、音频等模态与文本结合以实现类似能力。在此背景下,大音频语言模型(LALMs)需在不同于传统分类或生成任务的推理任务上进行评估。为此,我们提出了一个名为时间推理音频评估(TREA)的新数据集。对开源LALMs的基准测试显示,它们在TREA数据集的任务中始终落后于人类表现。评估过程中,我们还提出了一种不确定性度量,通过计算模型对语义相同输入扰动的不变性来衡量。分析表明,准确率与不确定性指标并非必然相关,这提示在高风险应用中需要更全面的LALMs评估体系。

原文摘要 · Abstract (English)

The popular success of text-based large language models (LLM) has streamlined the attention of the multimodal community to combine other modalities like vision and audio along with text to achieve similar multimodal capabilities. In this quest, large audio language models (LALMs) have to be evaluated on reasoning related tasks which are different from traditional classification or generation tasks. Towards this goal, we propose a novel dataset called temporal reasoning evaluation of audio (TREA). We benchmark open-source LALMs and observe that they are consistently behind human capabilities on the tasks in the TREA dataset. While evaluating LALMs, we also propose an uncertainty metric, which computes the invariance of the model to semantically identical perturbations of the input. Our analysis shows that the accuracy and uncertainty metrics are not necessarily correlated and thus, points to a need for wholesome evaluation of LALMs for high-stakes applications.

音频模型时间推理评估基准不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。