GMLv2用新损失函数提升音频质量预测精度
Enhanced Generative Machine Listener
- 采用贝塔分布损失建模听感评分,更贴合主观评价分布
- 在多种编码格式和内容上相关性超越PEAQ、ViSQOL等主流指标
- 适合音频编码研究者快速自动化评估音质
我们提出GMLv2,一种基于参考的模型,用于预测主观音频质量(以MUSHRA评分为标准)。GMLv2引入基于贝塔分布的损失函数来建模听者评分,并整合了额外的神经音频编码主观数据集,以增强其泛化能力和适用范围。在多样测试集上的大量评估表明,所提出的GMLv2在与主观评分的相关性以及跨多种内容类型和编码配置下可靠预测评分方面,均持续优于广泛使用的度量指标(如PEAQ和ViSQOL)。因此,GMLv2提供了一个可扩展且自动化的感知音频质量评估框架,有望加速现代音频编码技术的研究与开发。
原文摘要 · Abstract (English)
We present GMLv2, a reference-based model designed for the prediction of subjective audio quality as measured by MUSHRA scores. GMLv2 introduces a Beta distribution-based loss to model the listener ratings and incorporates additional neural audio coding (NAC) subjective datasets to extend its generalization and applicability. Extensive evaluations on diverse testset demonstrate that proposed GMLv2 consistently outperforms widely used metrics, such as PEAQ and ViSQOL, both in terms of correlation with subjective scores and in reliably predicting these scores across diverse content types and codec configurations. Consequently, GMLv2 offers a scalable and automated framework for perceptual audio quality evaluation, poised to accelerate research and development in modern audio coding technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。