通过评测挑战赛验证文本转音频生成效果,兼顾客观指标与人耳感知。
Challenge on Sound Scene Synthesis: Evaluating Text-to-Audio Generation
- 采用弗雷歇音频距离与人工评估结合的评测协议。
- 大模型整体表现更好,轻量级模型也展现潜力。
- 评测结果可指导未来文本转音频系统设计。
尽管神经文本到音频生成取得显著进展,可控性与评估仍存挑战。本文通过2024年声景检测与事件分类(DCASE 2024)中的声音场景合成挑战赛,提出一种结合客观指标(弗雷歇音频距离,FAD)与感知评估的评测协议,并使用结构化提示格式实现多样化描述词与高效评估。分析显示不同声音类别与模型架构间性能差异明显,大模型普遍表现更优,但创新的轻量级方法亦具前景。客观指标与人工评分高度相关,验证了评测方法的有效性。文章从音质、可控性及架构设计角度讨论成果,为未来研究提供方向。
原文摘要 · Abstract (English)
Despite significant advancements in neural text-to-audio generation, challenges persist in controllability and evaluation. This paper addresses these issues through the Sound Scene Synthesis challenge held as part of the Detection and Classification of Acoustic Scenes and Events 2024. We present an evaluation protocol combining objective metric, namely Fréchet Audio Distance, with perceptual assessments, utilizing a structured prompt format to enable diverse captions and effective evaluation. Our analysis reveals varying performance across sound categories and model architectures, with larger models generally excelling but innovative lightweight approaches also showing promise. The strong correlation between objective metrics and human ratings validates our evaluation approach. We discuss outcomes in terms of audio quality, controllability, and architectural considerations for text-to-audio synthesizers, providing direction for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。