arXiv:2507.04094eess.AScs.AI2025-07中稿 · ASRU Audio MOS 202…被引 2

多领域音频质量评估新模型,可细分四维度评分

MMMOS: Multi-domain Multi-axis Audio Quality Assessment

  • 融合三个预训练编码器的帧级特征,分维度评估音频质量
  • 在32项指标中17项进入前三,复杂度评估排名第一
  • 适合语音、音乐和环境音的多场景质量分析

精准的音频质量评估对音频生成、检索与增强系统的发展与评价至关重要。现有非侵入式评估模型仅针对语音输出单一平均意见分(MOS),合并了多种感知因素且泛化能力差。我们提出MMMOS,一种无参考的多领域音频质量评估系统,可在语音、音乐和环境音中估计四个正交维度:制作质量、制作复杂度、内容愉悦度和内容实用性。MMMOS融合来自三个预训练编码器(WavLM、MuQ、M2D)的帧级嵌入,并评估三种聚合策略与四种损失函数。通过集成表现最优的八组模型,MMMOS在均方误差上降低20-30%,肯德尔τ相关系数提升4-5%;在八项指标中的六项生产复杂度评估中位列第一,在32项挑战指标中17项排名前三。

原文摘要 · Abstract (English)

Accurate audio quality estimation is essential for developing and evaluating audio generation, retrieval, and enhancement systems. Existing non-intrusive assessment models predict a single Mean Opinion Score (MOS) for speech, merging diverse perceptual factors and failing to generalize beyond speech. We propose MMMOS, a no-reference, multi-domain audio quality assessment system that estimates four orthogonal axes: Production Quality, Production Complexity, Content Enjoyment, and Content Usefulness across speech, music, and environmental sounds. MMMOS fuses frame-level embeddings from three pretrained encoders (WavLM, MuQ, and M2D) and evaluates three aggregation strategies with four loss functions. By ensembling the top eight models, MMMOS shows a 20-30% reduction in mean squared error and a 4-5% increase in Kendall's τ versus baseline, gains first place in six of eight Production Complexity metrics, and ranks among the top three on 17 of 32 challenge metrics.

音频评估多维度无参考多领域

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。