arXiv:2508.08957cs.SDcs.AI2025-08中稿 · IEEE ASRU 2025被引 4

针对语音生成评估中主观感知差异问题,提出自适应排序优化方法提升评分准确率。

QAMRO: Quality-aware Adaptive Margin Ranking Optimization for Human-aligned Assessment of Audio Generation Systems

  • 基于多视角回归目标构建自适应边缘排序框架,捕捉感知差异
  • 在AudioMOS Challenge 2025数据集上显著优于基线模型
  • 适用于文本到音乐/语音等音频生成系统的质量评估

评估语音生成系统(包括文本到音乐、文本到语音、文本到音频)仍具挑战性,因其涉及主观且多维度的人类感知。现有方法将平均意见得分(MOS)预测视为回归问题,但标准回归损失忽略了感知判断的相对性。为此,我们提出QAMRO——一种质量感知的自适应边缘排序优化框架,无缝整合来自不同视角的回归目标,旨在凸显感知差异并优先保证评分准确性。该框架利用CLAP和Audiobox-Aesthetics等预训练音视频模型,仅在官方AudioMOS Challenge 2025数据集上进行训练,在所有维度上均表现出与人类评价更强的一致性,显著优于稳健基线模型。

原文摘要 · Abstract (English)

Evaluating audio generation systems, including text-to-music (TTM), text-to-speech (TTS), and text-to-audio (TTA), remains challenging due to the subjective and multi-dimensional nature of human perception. Existing methods treat mean opinion score (MOS) prediction as a regression problem, but standard regression losses overlook the relativity of perceptual judgments. To address this limitation, we introduce QAMRO, a novel Quality-aware Adaptive Margin Ranking Optimization framework that seamlessly integrates regression objectives from different perspectives, aiming to highlight perceptual differences and prioritize accurate ratings. Our framework leverages pre-trained audio-text models such as CLAP and Audiobox-Aesthetics, and is trained exclusively on the official AudioMOS Challenge 2025 dataset. It demonstrates superior alignment with human evaluations across all dimensions, significantly outperforming robust baseline models.

音频生成质量评估感知排序

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。