新基准BRACE揭示音频描述评估的局限性,挑战主流方法可靠性。
BRACE: A Benchmark for Robust Audio Caption Quality Evaluation
- 构建双子集基准,检测细粒度描述匹配与细微幻觉内容
- 最佳模型仅达70.01 F1,暴露现有评估与模型能力瓶颈
- 适合研究音频语言模型与评估指标的学者参考
自动音频描述对音频理解至关重要,广泛应用于无障碍和内容索引。然而,在缺乏高质量参考描述的无参考设置下,评估音频描述质量仍是重大挑战。尽管CLAPScore目前是最常用的无参考音频描述评估指标(ACEM),其在多样条件下的鲁棒性尚未系统验证。为此,我们提出BRACE,一个用于无参考环境下评估音频描述对齐质量的新基准。BRACE主要针对评估ACEM,也可扩展至衡量大音频语言模型(LALM)的模态对齐能力。该基准包含两个子集:BRACE-Main用于细粒度描述比较,BRACE-Hallucination用于检测细微幻觉内容。数据集通过高质量过滤、基于LLM的破坏生成和人工标注构建。鉴于CLAPScore被广泛采用且LALM在音频-语言任务中日益普及,我们使用BRACE评估多种CLAP模型变体和多个LALM。结果显示,即使表现最佳的基于CLAP的ACEM在BRACE-Main上也仅获70.01 F1得分,而最佳LALM仅达63.19。该结果揭示了CLAP模型与LALM的固有局限,为未来研究指明方向。
原文摘要 · Abstract (English)
Automatic audio captioning is essential for audio understanding, enabling applications such as accessibility and content indexing. However, evaluating the quality of audio captions remains a major challenge, especially in reference-free settings where high-quality ground-truth captions are unavailable. While CLAPScore is currently the most widely used reference-free Audio Caption Evaluation Metric(ACEM), its robustness under diverse conditions has not been systematically validated. To address this gap, we introduce BRACE, a new benchmark designed to evaluate audio caption alignment quality in a reference-free setting. BRACE is primarily designed for assessing ACEMs, and can also be extended to measure the modality alignment abilities of Large Audio Language Model(LALM). BRACE consists of two sub-benchmarks: BRACE-Main for fine-grained caption comparison and BRACE-Hallucination for detecting subtle hallucinated content. We construct these datasets through high-quality filtering, LLM-based corruption, and human annotation. Given the widespread adoption of CLAPScore as a reference-free ACEM and the increasing application of LALMs in audio-language tasks, we evaluate both approaches using the BRACE benchmark, testing CLAPScore across various CLAP model variants and assessing multiple LALMs. Notably, even the best-performing CLAP-based ACEM achieves only a 70.01 F1-score on the BRACE-Main benchmark, while the best LALM reaches just 63.19. By revealing the limitations of CLAP models and LALMs, our BRACE benchmark offers valuable insights into the direction of future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。