首个面向长音频理解的综合评测基准,揭示大模型在长音频上性能骤降的关键问题。
ChronosAudio: A Comprehensive Long-Audio Benchmark for Evaluating Audio-Large Language Models
- 构建覆盖6类任务的长音频评测集,含超200小时音频数据。
- 发现模型在长音频上性能下降超90%,注意力机制随序列变长严重发散。
- 现有缓解方法最多恢复50%性能,提示需新方法实现文档级音频推理。
尽管音频大语言模型(ALLMs)取得了显著进展,其对长音频的理解能力仍处于未探索状态。现有大量基准多聚焦于短时音频片段,缺乏对长音频理解的有效评估标准。本文提出ChronosAudio,首个专为长音频理解设计的多任务评测基准,涵盖六类主要任务,包含36,000个测试实例,总音频时长超过200小时,并按短、中、长三类分层,全面评估模型在不同长度上的泛化能力。在16个前沿模型上的实验揭示三个关键发现:1. 长上下文崩溃现象显著——从短到长的过渡导致特定任务性能下降超90%;2. 结构性注意力稀释:注意力机制在长序列中难以维持时间局部性,出现显著扩散;3. 缓解策略效果有限——当前方法仅能恢复约50%性能。这些结果揭示了长音频理解中的重大挑战,凸显了实现稳健的文档级音频推理的紧迫需求。
原文摘要 · Abstract (English)
Although Audio Large Language Models (ALLMs) have witnessed substantial advancements, their long audio understanding capabilities remain unexplored. A plethora of benchmarks have been proposed for general audio tasks, they predominantly focus on short-form clips, leaving without a consensus on evaluating ALLMs over extended durations. This paper proposes ChronosAudio, the first multi-task benchmark tailored for long-audio understanding in ALLMs. It encompasses six major task categories and comprises 36,000 test instances totaling over 200 hours audio, stratified into short, middle, and long-form categories to comprehensively evaluate length generalization. Extensive experiments on 16 state-of-the-art models using ChronosAudio yield three critical findings: 1.Precipitous Long-Context Collapse: ALLMs exhibit a severe inability to sustain performance, with the transition from short to long contexts triggering a staggering performance degradation of over 90% in specific tasks. 2.Structural Attention Dilution: Performance degradation stems from a fundamental failure in maintaining temporal locality; attention mechanisms suffer from significant diffusion in later sequences. 3.Restorative Ceiling of Mitigation: Current strategies only offer 50% recovery. These findings reveal significant challenges in long-audio, underscoring the urgent need for approaches to achieve robust, document-level audio reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。