提出针对性域策略,让音频质量评估更准、更通用。
Robust Generative Audio Quality Assessment: Disentangling Quality from Spurious Correlations
- 按质量维度定制领域划分,分离真实质量与虚假关联
- 在未见生成场景下相关性提升,超越现有方法
- 适合需要可靠音频评估的AI生成内容研究者
AI生成内容(AIGC)的快速发展催生了对感知质量评估指标的迫切需求。然而,自动均值意见分数(MOS)预测模型常因数据稀缺而学习到虚假相关性——如特定数据集的声学特征——而非泛化的质量特征。为此,我们采用领域对抗训练(DAT)来剥离真实质量感知与干扰因素。不同于依赖静态领域先验的先前工作,我们系统研究了从显式元数据标签到隐式数据驱动聚类等多种领域定义策略。结果表明,并不存在“通用适用”的领域定义;最优策略高度依赖于所评估的MOS具体维度。实验显示,针对不同维度采用特定领域策略可有效缓解声学偏差,在人类评分相关性上显著提升,并在未见过的生成场景中实现更优泛化性能。
原文摘要 · Abstract (English)
The rapid proliferation of AI-Generated Content (AIGC) has necessitated robust metrics for perceptual quality assessment. However, automatic Mean Opinion Score (MOS) prediction models are often compromised by data scarcity, predisposing them to learn spurious correlations-- such as dataset-specific acoustic signatures-- rather than generalized quality features. To address this, we leverage domain adversarial training (DAT) to disentangle true quality perception from these nuisance factors. Unlike prior works that rely on static domain priors, we systematically investigate domain definition strategies ranging from explicit metadata-driven labels to implicit data-driven clusters. Our findings reveal that there is no "one-size-fits-all" domain definition; instead, the optimal strategy is highly dependent on the specific MOS aspect being evaluated. Experimental results demonstrate that our aspect-specific domain strategy effectively mitigates acoustic biases, significantly improving correlation with human ratings and achieving superior generalization on unseen generative scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。