开源工具箱让音频大模型评估更高效、更全面,支持长对话分析。
AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs
- 通过批量优化和并行执行,提速最高达151%。
- 支持多轮对话,可研究模型跨轮次上下文理解能力。
- 适合需要系统评估音频大模型的研究者使用。
大型音频语言模型(LALMs)发展迅速,但评估仍面临效率低、标准不一的挑战,限制了公平比较与系统性评估。现有框架存在三大缺陷:(1)处理流程缓慢,阻碍大规模研究;(2)缺乏多轮对话支持,无法解答跨轮次上下文整合与性能演化等核心问题;(3)缺少统一且可扩展的评估体系,难以跟上模型与音频基准的快速迭代。为此,我们提出AU-Harness,一个高效且全面的LALM评估框架。该系统通过优化批处理与并行执行,相较现有工具提速最高达151%,使此前被认为不切实际的大规模评估成为可能。提供标准化提示协议与灵活配置,保障跨场景公平对比。AU-Harness支持多轮对话动态分析,揭示真实音频推理能力,既提供实用评估工具,也揭示模型局限,推动系统性LALM研发。
原文摘要 · Abstract (English)
Large Audio Language Models (LALMs) are rapidly advancing, but evaluating them remains challenging due to inefficient and non-standardized toolkits that limit fair comparison and systematic assessment. Existing evaluation frameworks exhibit three critical limitations: (1) slow and inefficient processing pipeline that bottlenecks large-scale studies, (2) inadequate multi-turn dialogue support, leaving fundamental questions about cross-turn context integration and performance dynamics over extended conversations in LALMs unanswered; and (3) the absence of unified and scalable evaluation framework capable of keeping pace with the rapid growth of both LALMs and audio benchmarks. To address these issues, we introduce AU-Harness, an efficient and comprehensive evaluation framework for LALMs. Our system achieves a speedup of up to 151% over existing evaluation toolkits through optimized batch processing and parallel execution, enabling large-scale evaluations previously considered impractical. We provide standardized prompting protocols and flexible configurations for fair model comparison across diverse scenarios. AU-Harness unlocks a range of in-depth analyses difficult to conduct without a unified foundation, including multi-turn dialogue dynamics, enabling the study of true audio reasoning capabilities in existing LALMs. AU-Harness provides both practical evaluation tools and insights into model limitations, advancing systematic LALM development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。