构建跨声学领域的音频描述评估基准,测试大模型表现。
AudioCapBench: Quick Evaluation on Audio Captioning across Sound, Music, and Speech
- 构建涵盖环境音、音乐、语音三类的1000样本评估集。
- 谷歌Gemini模型整体表现最优,但开放模型幻觉更少。
- 模型在语音上表现最好,音乐上最差,评估覆盖准确性等三维度。
我们提出AudioCapBench,一个用于评估大型多模态模型音频描述能力的基准。该基准覆盖环境音、音乐和语音三大音频领域,包含从已有数据集精选的1000个评估样本。我们使用参考指标(METEOR、BLEU、ROUGE-L)及基于LLM的评判框架,从准确性、完整性与幻觉三个维度评估13个模型(来自OpenAI和Google Gemini)。结果表明,Gemini模型整体表现优于OpenAI模型,其中Gemini 3 Pro得分最高(6.00/10),而OpenAI模型幻觉率更低。所有模型在语音描述任务中表现最佳,在音乐描述中最差。我们开源了基准与评估代码,支持可复现的音频理解研究。
原文摘要 · Abstract (English)
We introduce AudioCapBench, a benchmark for evaluating audio captioning capabilities of large multimodal models. \method covers three distinct audio domains, including environmental sound, music, and speech, with 1,000 curated evaluation samples drawn from established datasets. We evaluate 13 models across two providers (OpenAI, Google Gemini) using both reference-based metrics (METEOR, BLEU, ROUGE-L) and an LLM-as-Judge framework that scores predictions on three orthogonal dimensions: \textit{accuracy} (semantic correctness), \textit{completeness} (coverage of reference content), and \textit{hallucination} (absence of fabricated content). Our results reveal that Gemini models generally outperform OpenAI models on overall captioning quality, with Gemini~3~Pro achieving the highest overall score (6.00/10), while OpenAI models exhibit lower hallucination rates. All models perform best on speech captioning and worst on music captioning. We release the benchmark as well as evaluation code to facilitate reproducible audio understanding research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。