arXiv:2502.05139cs.SDcs.LG2025-02被引 193

提出统一音频美学评估方法,无需人工打分即可自动判断语音音乐质量。

Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

  • 将人类听感分解为四个维度,制定新标注规范
  • 构建无参考的单条音频评分模型,性能媲美人工评分
  • 适合生成模型评估与数据清洗,开源可用

音频美学的量化仍是音频处理中的难题,主要因其主观性受人耳感知与文化背景影响。传统方法依赖人工听评,存在不一致且成本高。本文针对自动化音频美学评估的需求,提出新的标注指南,将人类听觉视角分解为四个独立维度,并开发训练无参考、逐项预测的模型,实现更细致的音频质量评估。模型在与人工平均意见分(MOS)及现有方法对比中表现相当或更优。本研究不仅推动音频美学领域发展,还开放源代码与预训练模型,供后续研究与基准测试使用。

原文摘要 · Abstract (English)

The quantification of audio aesthetics remains a complex challenge in audio processing, primarily due to its subjective nature, which is influenced by human perception and cultural context. Traditional methods often depend on human listeners for evaluation, leading to inconsistencies and high resource demands. This paper addresses the growing need for automated systems capable of predicting audio aesthetics without human intervention. Such systems are crucial for applications like data filtering, pseudo-labeling large datasets, and evaluating generative audio models, especially as these models become more sophisticated. In this work, we introduce a novel approach to audio aesthetic evaluation by proposing new annotation guidelines that decompose human listening perspectives into four distinct axes. We develop and train no-reference, per-item prediction models that offer a more nuanced assessment of audio quality. Our models are evaluated against human mean opinion scores (MOS) and existing methods, demonstrating comparable or superior performance. This research not only advances the field of audio aesthetics but also provides open-source models and datasets to facilitate future work and benchmarking. We release our code and pre-trained model at: https://github.com/facebookresearch/audiobox-aesthetics

音频评估生成模型无参考开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。