arXiv:2509.19219eess.AS2025-09被引 2

提出新型语音评估方法,兼顾敏感度与可扩展性。

MUSHRA-1S: A scalable and sensitive test approach for evaluating top-tier speech processing systems

  • 单刺激评估,固定参考对比,逐系统评分
  • 高保真场景下仍能分辨细微差异,不饱和
  • 适合顶级语音系统评测,减少评分偏差

评估顶尖语音处理系统需要兼具可扩展性与敏感性的评测方法,以检测细微但不可接受的失真。标准MUSHRA虽敏感但难以扩展,ACR可扩展却失去敏感性且在高质量场景下趋于饱和。为此,本文提出MUSHRA 1S,一种单刺激评估方法,每次仅对一个系统进行评分,以固定锚点和参考为基准。实验表明,MUSHRA 1S在高保真区域的表现比ACR更接近标准MUSHRA,能有效识别特定失真,并通过固定上下文减少范围均衡偏差。该方法结合了MUSHRA的高敏感性与ACR的可扩展性,为顶级语音系统的基准测试提供可靠方案。

原文摘要 · Abstract (English)

Evaluating state-of-the-art speech systems necessitates scalable and sensitive evaluation methods to detect subtle but unacceptable artifacts. Standard MUSHRA is sensitive but lacks scalability, while ACR scales well but loses sensitivity and saturates at a high quality. To address this, we introduce MUSHRA 1S, a single-stimulus variant that rates one system at a time against a fixed anchor and reference. Across our experiments, MUSHRA 1S matches standard MUSHRA more closely than ACR, including in the high-quality regime, where ACR saturates. MUSHRA 1S also effectively identifies specific deviations and reduces range-equalizing biases by fixing context. Overall, MUSHRA 1S combines MUSHRA level sensitivity with ACR like scalability, making it a robust and scalable solution for benchmarking top-tier speech processing systems.

语音评估MUSHRA可扩展性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。