用众包方式测试生成式语音模型音质,更高效可靠。
Crowdsourcing MUSHRA Tests in the Age of Generative Speech Technologies: A Comparative Analysis of Subjective and Objective Testing Methods
- 用众包非专家替代专业听众做MUSHRA测评
- 发现传统客观指标低估生成式语音质量
- 适合语音模型开发者和评测人员参考
MUSHRA框架广泛用于检测音频质量的细微差异,但传统上依赖受控环境中专家听者,成本高且不适用于模型迭代。因此开发阶段常使用客观指标,后期再进行专家评估。然而,这些指标对生成式语音模型效果不佳。本文提出针对生成式语音编码器的众包MUSHRA测试方法,通过MTurk与Prolific平台对比专家数据,验证了测试重测信度与一致性。同时评估六种客观指标,发现传统指标普遍低估生成式模型表现。研究揭示平台特异性偏差,强调需采用编码器感知型指标,为语音编码器的可扩展感知测试提供指导。
原文摘要 · Abstract (English)
The MUSHRA framework is widely used for detecting subtle audio quality differences but traditionally relies on expert listeners in controlled environments, making it costly and impractical for model development. As a result, objective metrics are often used during development, with expert evaluations conducted later. While effective for traditional DSP codecs, these metrics often fail to reliably evaluate generative models. This paper proposes adaptations for conducting MUSHRA tests with non-expert, crowdsourced listeners, focusing on generative speech codecs. We validate our approach by comparing results from MTurk and Prolific crowdsourcing platforms with expert listener data, assessing test-retest reliability and alignment. Additionally, we evaluate six objective metrics, showing that traditional metrics undervalue generative models. Our findings reveal platform-specific biases and emphasize codec-aware metrics, offering guidance for scalable perceptual testing of speech codecs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。