arXiv:2603.19615cs.SDcs.AI2026-03

用大模型推理校准音频描述评分,更准识别语法错误和细节漏洞。

CAF-Score: Calibrating CLAP with LALMs for Reference-free Audio Captioning Evaluation

  • 融合对比嵌入与大模型推理,实现粗粒度语义与细粒度理解的结合。
  • 在BRACE数据集上与人工评分相关性最高,挑战场景下超越有参考基线。
  • 适合无参考文本的音频描述评估,尤其关注语法和细节真实性的研究者。

大型音频语言模型(LALMs)虽推动了音频描述的发展,但可靠评估仍具挑战。基于参考的指标成本高且难以衡量声学保真度,而基于对比语言-音频预训练(CLAP)的方法常忽略语法错误和细微细节。本文提出CAF-Score,一种无需参考文本的评估指标,通过将CLAP的粗粒度语义对齐与LALMs的细粒度理解及语法敏感性相结合,有效检测语法不一致与微小幻觉。在BRACE基准上的实验表明,该方法与人类判断的相关性最高,甚至在困难场景中优于基于参考的基线。代码与结果见https://github.com/inseong00/CAF-Score。

原文摘要 · Abstract (English)

While Large Audio-Language Models (LALMs) have advanced audio captioning, robust evaluation remains difficult. Reference-based metrics are expensive and often fail to assess acoustic fidelity, while Contrastive Language-Audio Pretraining (CLAP)-based approaches frequently overlook syntactic errors and fine-grained details. We propose CAF-Score, a reference-free metric that calibrates CLAP's coarse-grained semantic alignment with the fine-grained comprehension and syntactic awareness of LALMs. By combining contrastive audio-text embeddings with LALM reasoning, CAF-Score effectively detects syntactic inconsistencies and subtle hallucinations. Experiments on the BRACE benchmark demonstrate that our approach achieves the highest correlation with human judgments, even outperforming reference-based baselines in challenging scenarios. These results highlight the efficacy of CAF-Score for reference-free audio captioning evaluation. Code and results are available at https://github.com/inseong00/CAF-Score.

音频生成评估指标大模型无参考

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。