arXiv:2510.07293cs.SDcs.AI2025-10被引 23

构建长音频理解与效率评估新基准,推动语音大模型突破时序瓶颈

AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficiency in Audio LLMs

  • 设计涵盖90-300秒长音频的多任务评测体系,支持超7500个音频令牌输入
  • 发现现有模型在长音频下性能显著下降,推理效率与记忆机制成关键短板
  • 适合研究音频大模型、高效推理和多跳推理的学者使用

长音频处理是语音大模型(LALMs)的核心挑战,其注意力机制存在二次方复杂度($O(N^2)$)问题,且难以建模长时程依赖。现有音频基准多基于短片段,无法真实评估模型在长上下文下的表现。为此,我们提出AudioMarathon,一个全面评估长音频理解与推理效率的基准。该基准包含三方面:输入音频长度为90.0至300.0秒,对应2,250至7,500个音频标记;覆盖语音、声音和音乐全领域;要求多跳推理的复杂任务。我们评估了前沿的LALMs,观察到随着音频长度增加性能明显下降。同时研究了加速技术,分析了令牌剪枝与键值缓存剔除的权衡。结果揭示当前模型间存在显著差距,凸显对更优时序推理与内存高效架构的需求。我们认为AudioMarathon将推动音频与多模态研究向更复杂音频任务迈进。

原文摘要 · Abstract (English)

Processing long-form audio is a major challenge for Large Audio Language models (LALMs). These models struggle with the quadratic cost of attention ($O(N^2)$) and with modeling long-range temporal dependencies. Existing audio benchmarks are built mostly from short clips and do not evaluate models in realistic long context settings. To address this gap, we introduce AudioMarathon, a benchmark designed to evaluate both understanding and inference efficiency on long-form audio. AudioMarathon provides a diverse set of tasks built upon three pillars: long-context audio inputs with durations ranging from 90.0 to 300.0 seconds, which correspond to encoded sequences of 2,250 to 7,500 audio tokens, respectively, full domain coverage across speech, sound, and music, and complex reasoning that requires multi-hop inference. We evaluate state-of-the-art LALMs and observe clear performance drops as audio length grows. We also study acceleration techniques and analyze the trade-offs of token pruning and KV cache eviction. The results show large gaps across current LALMs and highlight the need for better temporal reasoning and memory-efficient architectures. We believe AudioMarathon will drive the audio and multimodal research community to develop more advanced audio understanding models capable of solving complex audio tasks.

长音频语音大模型推理效率多跳推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。