评测声音场景合成系统,揭示当前技术优劣与改进方向。
Sound Scene Synthesis at the DCASE 2024 Challenge
- 构建标准化评估框架,融合客观与主观评价指标。
- 4个参赛系统经FAD和人耳评分测试,结果可比。
- 揭示生成多样性与真实感的平衡难题,适合音频生成研究者参考。
本文介绍DCASE 2024挑战赛中的第7项任务:声音场景合成。近年来,声音合成与生成模型的发展推动了逼真且多样音频内容的生成。我们提出一个标准化评估框架,用于对比不同声音场景合成系统,涵盖客观与主观评价指标。本次挑战共吸引4个提交方案,采用弗雷谢音频距离(FAD)与人工感知评分进行评估。分析揭示了当前声音场景合成系统在真实性和多样性方面的实际能力与局限,同时也指出了该快速演进领域未来改进的关键方向。
原文摘要 · Abstract (English)
This paper presents Task 7 at the DCASE 2024 Challenge: sound scene synthesis. Recent advances in sound synthesis and generative models have enabled the creation of realistic and diverse audio content. We introduce a standardized evaluation framework for comparing different sound scene synthesis systems, incorporating both objective and subjective metrics. The challenge attracted four submissions, which are evaluated using the Fréchet Audio Distance (FAD) and human perceptual ratings. Our analysis reveals significant insights into the current capabilities and limitations of sound scene synthesis systems, while also highlighting areas for future improvement in this rapidly evolving field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。