构建大规模音频生成评估数据集,支持多维度自动评价。
AudioEval: Automatic Dual-Perspective and Multi-Dimensional Evaluation of Text-to-Audio-Generation
- 构建包含4200段音频的双视角多维度评估数据集
- 发现不同模型在质量与文本对齐上表现差异显著
- 适合研究音频生成评估与评测基准构建者
文本到音频(TTA)生成快速发展,但评估仍具挑战性,因人工听感测试成本高,现有自动指标仅捕捉感知质量的有限方面。我们提出AudioEval,一个大规模的TTA评估数据集,包含24个系统生成的4,200段音频(总计11.7小时),以及来自专家和非专家的126,000条评分,涵盖五维:愉悦度、实用性、复杂度、质量与文本对齐。基于AudioEval,我们对多种自动评估器进行基准测试,比较不同模型族在视角与维度上的差异。我们还提出Qwen-DisQA作为强基线:它联合处理提示与生成音频,预测两组标注者的多维评分,通过分布预测建模标注者分歧,表现优异。我们将公开AudioEval以支持未来TTA评估研究。
原文摘要 · Abstract (English)
Text-to-audio (TTA) generation is advancing rapidly, but evaluation remains challenging because human listening studies are expensive and existing automatic metrics capture only limited aspects of perceptual quality. We introduce AudioEval, a large-scale TTA evaluation dataset with 4,200 generated audio samples (11.7 hours) from 24 systems and 126,000 ratings collected from both experts and non-experts across five dimensions: enjoyment, usefulness, complexity, quality, and text alignment. Using AudioEval, we benchmark diverse automatic evaluators to compare perspective- and dimension-level differences across model families. We also propose Qwen-DisQA as a strong reference baseline: it jointly processes prompts and generated audio to predict multi-dimensional ratings for both annotator groups, modeling rater disagreement via distributional prediction and achieving strong performance. We will release AudioEval to support future research in TTA evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。