构建首个开放音频对话理解基准,评估大模型在多语言、歧义场景下的对话能力。
Benchmarking Open-ended Audio Dialogue Understanding for Large Audio-Language Models
- 设计4个数据集,覆盖3类场景、12项技能、9种语言和4类歧义处理。
- 超过2万条音频对话,揭示现有模型在数学符号、角色扮演和语调歧义上表现差。
- 适合研究语音交互、多模态理解或音视频大模型的开发者与研究者。
大型音频语言模型(LALMs)如GPT-4o最近实现了语音对话能力,支持与人类进行直接口语交流。这一进展拓展了其在各类音频对话应用场景中的潜力。然而,当前尚缺乏针对开放式音频对话理解的全面评估基准。为此,我们提出音频对话理解基准(ADU-Bench),包含4个数据集,用于评估LALMs在3类通用场景、12项技能、9种多语言及4类歧义处理上的表现。特别地,我们首次引入对音频对话中歧义的评估,即同一字面意思因语调不同而表达不同意图的情况,例如“真的吗?!”的不同语调。总体而言,ADU-Bench涵盖超过20,000条开放式音频对话,用于评估LALMs。通过对16个LALMs的广泛实验分析发现,现有模型在数学符号与公式理解、角色扮演行为识别、多语言理解以及由语调、停顿位置和同音词等语音元素引发的对话歧义处理方面仍存在显著困难。该基准可在 https://adu-bench.github.io/ 获取。
原文摘要 · Abstract (English)
Large Audio-Language Models (LALMs), such as GPT-4o, have recently unlocked audio dialogue capabilities, enabling direct spoken exchanges with humans. The potential of LALMs broadens their applicability across a wide range of practical scenarios supported by audio dialogues. However, given these advancements, a comprehensive benchmark to evaluate the performance of LALMs in the open-ended audio dialogue understanding remains absent currently. To address this gap, we propose an Audio Dialogue Understanding Benchmark (ADU-Bench), which consists of 4 benchmark datasets. They assess the open-ended audio dialogue ability for LALMs in 3 general scenarios, 12 skills, 9 multilingual languages, and 4 categories of ambiguity handling. Notably, we firstly propose the evaluation of ambiguity handling in audio dialogues that expresses different intentions beyond the same literal meaning of sentences, e.g., "Really!?" with different intonations. In summary, ADU-Bench includes over 20,000 open-ended audio dialogues for the assessment of LALMs. Through extensive experiments on 16 LALMs, our analysis reveals that existing LALMs struggle with mathematical symbols and formulas, understanding human behavior such as roleplay, comprehending multiple languages, and handling audio dialogue ambiguities from different phonetic elements, such as intonations, pause positions, and homophones. The benchmark is available at https://adu-bench.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。