arXiv:2510.12326eess.AS2025-10

用弱监督学习让音乐大模型感知音频失真,效果超越现有评测工具。

DeePAQ: A Perceptual Audio Quality Metric Based On Foundational Models and Weakly Supervised Learning

  • 基于音乐大模型MERT与度量学习构建失真感知嵌入空间
  • 在音频编码和源分离任务中均优于现有客观指标
  • 适用于未见过的失真类型,适合音频质量评估研究者

本文提出一种基于深度学习的通用音频质量感知评估方法DeePAQ。该方法结合度量学习与音乐基础模型MERT,利用代理标签引导,构建捕捉音频失真强度的嵌入空间。据我们所知,DeePAQ是首个在通用音频质量领域采用弱监督标签与度量学习,通过低秩适配(LoRA)微调音乐基础模型的方法,这一方向尚未被其他先进方法探索。我们在涵盖音频编码与源分离的听觉测试中对模型进行基准测试,结果表明,该方法在检测编码伪影方面表现更优,并能良好泛化至未见失真类型(如源分离),展现出更强的鲁棒性与适用性。

原文摘要 · Abstract (English)

This paper presents the Deep learning-based Perceptual Audio Quality metric (DeePAQ) for evaluating general audio quality. Our approach leverages metric learning together with the music foundation model MERT, guided by surrogate labels, to construct an embedding space that captures distortion intensity in general audio. To the best of our knowledge, DeePAQ is the first in the general audio quality domain to leverage weakly supervised labels and metric learning for fine-tuning a music foundation model with Low-Rank Adaptation (LoRA), a direction not yet explored by other state-of-the-art methods. We benchmark the proposed model against state-of-the-art objective audio quality metrics across listening tests spanning audio coding and source separation. Results show that our method surpasses existing metrics in detecting coding artifacts and generalizes well to unseen distortions such as source separation, highlighting its robustness and versatility.

音频质量大模型弱监督度量学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。