用音频大模型问答评估文本到音频生成的语义对齐,更准。
AQAScore: Evaluating Semantic Alignment in Text-to-Audio Generation via Audio Question Answering
- 用音频大模型回答特定问题判断语义对齐,不靠生成文本。
- 在多个评测中与人工判断相关性更高,能发现细微语义错误。
- 适合研究文本到音频生成、评估方法改进的学者使用。
尽管文本到音频生成在真实感和多样性上取得显著进展,但评估指标的发展仍滞后。现有方法如基于嵌入相似性的CLAPScore虽能衡量整体相关性,但在细粒度语义对齐和组合推理方面仍有不足。为此,我们提出AQAScore,一种不依赖主干模型的评估框架,利用音频感知大语言模型(ALLMs)的推理能力。AQAScore将评估任务重构为概率化语义验证:不依赖开放式文本生成,而是通过计算针对特定语义问题回答“是”的精确对数概率来估计对齐程度。我们在多个基准上进行了评估,包括人工标注的相关性、成对比较和组合推理任务。实验结果表明,AQAScore在与人工判断的相关性上持续优于基于相似性的度量方法和生成提示基线,有效捕捉细微语义不一致,并随底层ALLMs能力提升而增强。
原文摘要 · Abstract (English)
Although text-to-audio generation has made remarkable progress in realism and diversity, the development of evaluation metrics has not kept pace. Widely-adopted approaches, typically based on embedding similarity like CLAPScore, effectively measure general relevance but remain limited in fine-grained semantic alignment and compositional reasoning. To address this, we introduce AQAScore, a backbone-agnostic evaluation framework that leverages the reasoning capabilities of audio-aware large language models (ALLMs). AQAScore reformulates assessment as a probabilistic semantic verification task; rather than relying on open-ended text generation, it estimates alignment by computing the exact log-probability of a "Yes" answer to targeted semantic queries. We evaluate AQAScore across multiple benchmarks, including human-rated relevance, pairwise comparison, and compositional reasoning tasks. Experimental results show that AQAScore consistently achieves higher correlation with human judgments than similarity-based metrics and generative prompting baselines, showing its effectiveness in capturing subtle semantic inconsistencies and scaling with the capability of underlying ALLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。