arXiv:2605.27976cs.SD2026-05被引 3

构建首个小时级音频语言理解基准,揭示长时记忆瓶颈。

VoiceGiraffe: A Benchmark for Extreme Long-Context Audio-Language Understanding

论文配图:VoiceGiraffe: A Benchmark for Extreme Long-Context Audio-Language Understanding
图 1 · 摘自论文原文
  • 设计1500个真实场景三元组,分单跳感知与多跳推理双层级评测
  • 模型在长音频理解上远未达标,不同推理范式各有优劣
  • 揭示模型长时记忆弱于人类,尤其难追踪稀疏事件

尽管大音频语言模型(LALMs)在秒级至分钟级音频处理上取得显著进展,但对小时级音频的理解仍是根本性瓶颈。现有基准多依赖短片段或人为拼接,无法真实评估模型在播客、长篇演讲等场景下的长程信息理解能力。为此,我们提出VoiceGiraffe,一个新颖的基准,用于在多样真实场景、模态和语言下严格评估LALMs的长上下文能力。它包含1500个精心设计的三元组,采用双层级分类体系:单跳感知与多跳推理。我们评估了广泛开源与专有LALMs,并与人类表现对比。结果揭示三大发现:第一,VoiceGiraffe仍极具挑战性,尚未达到饱和;第二,无单一推理范式通用最优——端到端适合具备原生长时音频理解能力的模型,级联字幕聚合可稳定小模型在小时级音频下的表现,而引入外部大模型的增强级联虽助弱模型,却可能拖累强专有系统;第三,长时记忆持久性是关键瓶颈——模型更擅长关联关键因果线索,而人类则更擅长持续追踪稀疏事件。这些发现使VoiceGiraffe成为诊断长时音频理解能力的挑战性测试平台,凸显对具备持久记忆与鲁棒长程聚合能力的LALMs的迫切需求。

原文摘要 · Abstract (English)

While large audio language models (LALMs) have achieved remarkable progress in audio processing at the second- or minute-level scale, understanding hour-level audio remains a fundamental bottleneck. Existing benchmarks predominantly rely on short clips or artificially concatenated segments, failing to faithfully assess LALM capacity for long-range information comprehension in real-world scenarios such as podcasts and lengthy speeches. To address this gap, we introduce VoiceGiraffe, a novel benchmark designed to rigorously evaluate LALMs across diverse real-world scenarios, modalities, and languages under long-context settings. It comprises 1500 curated triplets structured into a dual-level taxonomy of single-hop perception and multi-hop reasoning. We evaluate a broad suite of open-source and proprietary LALMs against human performance. Results underscore three fundamental findings. First, VoiceGiraffe remains highly challenging and far from saturation. Second, we show that no single inference paradigm universally dominates. The E2E inference benefits models with native long-context audio understanding, cascaded caption aggregation stabilizes small models overwhelmed by hour-scale audio, and reasoning-enhanced cascading with external LLM helps weaker models but can bottleneck stronger proprietary systems. Third, we reveal long-range memory persistence as a key bottleneck. LALMs are better at answering questions that require connecting salient causal cues than those requiring sustained tracking of sparse events across long audio, whereas humans show the opposite pattern. These findings position VoiceGiraffe as a challenging and diagnostic testbed for long-form audio understanding, highlighting the need for LALMs with persistent memory and robust long-range aggregation.

音频理解长上下文基准评测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。