无需参考音频,可精准评估各类编码音频质量
RF-GML: Reference-Free Generative Machine Listener
- 基于先进模型迁移,不依赖原始参考信号生成评分
- 在多种内容与编码格式下准确预测主观评分
- 适合音频质量评估、监控等无需参考信号的场景
本文提出一种新型无参考(RF)音频质量度量方法——参考无生成式机器听者(RF-GML),用于评估48 kHz采样率下的单声道、立体声及双耳编码音频。RF-GML通过最小改动迁移自最先进的全参考生成式机器听者(GML)模型,具备生成任意数量模拟听感评分的能力。与现有无参考模型不同,RF-GML可在多样内容类型和编码方式下准确预测主观质量评分。大量实验表明,其在未编码音频评分及区分不同编码失真等级方面表现优异。该模型性能强、适用广,适用于各类无需参考信号的编码音频质量评估与监控任务。
原文摘要 · Abstract (English)
This paper introduces a novel reference-free (RF) audio quality metric called the RF-Generative Machine Listener (RF-GML), designed to evaluate coded mono, stereo, and binaural audio at a 48 kHz sample rate. RF-GML leverages transfer learning from a state-of-the-art full-reference (FR) Generative Machine Listener (GML) with minimal architectural modifications. The term "generative" refers to the model's ability to generate an arbitrary number of simulated listening scores. Unlike existing RF models, RF-GML accurately predicts subjective quality scores across diverse content types and codecs. Extensive evaluations demonstrate its superiority in rating unencoded audio and distinguishing different levels of coding artifacts. RF-GML's performance and versatility make it a valuable tool for coded audio quality assessment and monitoring in various applications, all without the need for a reference signal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。