arXiv:2607.27109cs.SDcs.AI2026-07

构建首个大规模多维度音频描述评估基准,提升生成内容的细节覆盖与可靠性。

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning

论文配图:MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning
图 1 · 摘自论文原文
  • 设计多维度评估框架,涵盖6类能力与15个评测维度
  • 包含5638段跨源音频数据,支持细粒度内容覆盖检测
  • 适合研究音频大模型生成质量与信息可靠性的团队使用

随着音频大语言模型(AudioLLMs)的发展,音频描述任务需从简短描述转向开放且细致的自由形式描述。现有评估多关注生成质量或任务性能,难以诊断信息覆盖度与描述可靠性。我们提出MMAC——一个大规模多维度音频描述评估基准。该基准包含来自20多个数据源的5,638段音频,覆盖6类能力与15个评估维度。针对模型生成的描述,MMAC检查其是否提及目标维度的相关信息,以及所述内容是否与参考标签一致。我们评估了代表性开源与专有AudioLLMs,结果揭示不同维度间存在显著差异,涵盖信息覆盖度与描述可靠性。我们将公开MMAC基准与评估代码。

原文摘要 · Abstract (English)

With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-grained free-form descriptions. Existing evaluations often focus on generation quality or task performance, making it difficult to diagnose information coverage and description reliability. We propose MMAC, a \textbf{M}assive \textbf{M}ulti-dimensional benchmark for \textbf{A}udio \textbf{C}aptioning. MMAC contains 5,638 audio clips from more than 20 data sources, covering 6 capability categories and 15 evaluation dimensions. Given a model-generated caption, MMAC checks whether it mentions relevant information in the target dimension and whether the mentioned content is consistent with the reference label. We evaluate representative open-source and proprietary AudioLLMs. Results show clear differences across evaluation dimensions, information coverage, and description reliability. We will release the MMAC benchmark and evaluation code.

音频描述评估基准大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。