构建首个覆盖多维度的文本转音频评估基准,全面检验模型性能与伦理风险。
TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models
- 设计涵盖7个维度的评估框架,融合自动化与人工生成的2999个多样化提示
- 采用超11.8万次人评标注,量化模型在准确性、鲁棒性、毒性等表现
- 开源完整工具链,适合研究者和开发者用于模型可信度评测
文本转音频(TTA)生成技术发展迅速,但现有评估方法仍局限于感知质量,忽视了鲁棒性、泛化能力及伦理问题。我们提出TTA-Bench,一个涵盖功能性能、可靠性与社会责任的综合性评估基准。该基准包含7个维度:准确性、鲁棒性、公平性、毒性等,涵盖2,999个通过自动化与人工方式生成的多样化提示。我们引入统一评估协议,结合客观指标与超过118,000人次的人工标注(来自专家与普通用户)。在该框架下对10个先进模型进行评测,揭示其优缺点。TTA-Bench为TTA系统提供了全新的全方位、负责任的评估标准。数据集与评估工具已开源,地址:https://nku-hlt.github.io/tta-bench/
原文摘要 · Abstract (English)
Text-to-Audio (TTA) generation has made rapid progress, but current evaluation methods remain narrow, focusing mainly on perceptual quality while overlooking robustness, generalization, and ethical concerns. We present TTA-Bench, a comprehensive benchmark for evaluating TTA models across functional performance, reliability, and social responsibility. It covers seven dimensions including accuracy, robustness, fairness, and toxicity, and includes 2,999 diverse prompts generated through automated and manual methods. We introduce a unified evaluation protocol that combines objective metrics with over 118,000 human annotations from both experts and general users. Ten state-of-the-art models are benchmarked under this framework, offering detailed insights into their strengths and limitations. TTA-Bench establishes a new standard for holistic and responsible evaluation of TTA systems. The dataset and evaluation tools are open-sourced at https://nku-hlt.github.io/tta-bench/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。