改进众包听感测试,提升语音编码评估的可靠性与成本效益。
Screening Matters: A Comparative Study of Conventional and Crowdsourced Listening Tests
- 通过锚点排序和评分跨度等后筛选方法优化众包测试
- 陷阱题与黄金标准题实现持续筛选,显著提升数据质量
- 提出低成本、低偏差的听感评估方案,适合工业级语音测试
主观评估仍是检验语音与音频编码技术最可靠的方式。众包听感测试虽具成本低、速度快的优势,但结果质量通常不如实验室控制环境下的传统测试。本文对比评估了经典与神经语音编解码器,采用P.808与P.800 DCR测试方法,并通过统计分析研究多种筛选策略的有效性。结果表明,结合锚点排序与评分跨度的后筛选方法,以及陷阱题、黄金标准题等持续筛选手段,可有效提升众包测试的数据可信度,增强对被测编码器评分的价值。基于此,本文提出一套适用于成本可控、流程简化且偏差低的听感评价筛选组合。
原文摘要 · Abstract (English)
Subjective evaluation remains the most reliable way of testing speech and audio coding techniques. Crowdsourcing the listening task is a cost-efficient and fast way of conducting this evaluation, but the quality of the results tends to be inferior to that of conventional listening tests done in the controlled environment of a laboratory. In this paper, classical and neural speech codecs are evaluated to compare P.808 against P.800 DCR tests. A statistical analysis is conducted to investigate the effectiveness of selected screening methods. The analysis shows that the crowdsourced evaluation can be improved by employing postscreening methods based on anchor ordering and rating span, and continuous screening methods like traps and gold standard questions, thus giving more value to the ratings obtained for the codecs under test. Based on these outcomes, a set of suitable screenings is proposed, for cost-effective, simplified, and bias-free enhancement of listening results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。