arXiv:2509.01336cs.SDeess.AS2025-09被引 13

首个合成音频主观质量评估挑战,推动自动评测技术发展

The AudioMOS Challenge 2025

  • 设立三赛道,分别评估文本到音乐、语音与音频的生成质量
  • 24支来自学界与产业的团队参与,结果均优于基线模型
  • 为语音与音乐生成系统提供标准化自动评估基准

本文是 AudioMOS Challenge 2025 的总结论文,该挑战是首个针对合成音频的自动主观质量预测竞赛。挑战包含三个赛道:第一赛道评估文本到音乐样本的整体质量与文本对齐度;第二赛道基于 Meta Audiobox Aesthetics 的四个维度,测试集涵盖文本到语音、文本到音频及文本到音乐样本;第三赛道聚焦不同采样率下的合成语音质量评估。共有24支来自学术界与工业界的独立团队参与,所有提交方案均在各项任务上实现对基线模型的改进。该挑战成果有望推动音频生成系统自动评估技术的发展与进步。

原文摘要 · Abstract (English)

This is the summary paper for the AudioMOS Challenge 2025, the very first challenge for automatic subjective quality prediction for synthetic audio. The challenge consists of three tracks. The first track aims to assess text-to-music samples in terms of overall quality and textual alignment. The second track is based on the four evaluation dimensions of Meta Audiobox Aesthetics, and the test set consists of text-to-speech, text-to-audio, and text-to-music samples. The third track focuses on synthetic speech quality assessment in different sampling rates. The challenge attracted 24 unique teams from both academia and industry, and improvements over the baselines were confirmed. The outcome of this challenge is expected to facilitate development and progress in the field of automatic evaluation for audio generation systems.

音频评估主观质量生成评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。