提出可精确定位语音合成质量异常帧的评估方法
Towards Frame-level Quality Predictions of Synthetic Speech
- 设计分块处理机制,避免局部失真影响邻近帧评分
- 在人工添加失真数据上,模型定位精度超过众包人类标注
- 为语音合成系统提供可解释的质量评估新思路
尽管自动主观语音质量评估已取得显著进展,但能否实现帧级质量评估仍是一个开放问题。这将极大提升语音合成系统评估的可解释性。本文首次探索该目标,识别现有质量预测器在帧级预测中的缺陷,并定义了帧级预测应满足的标准。提出分块处理策略,避免局部失真对相邻帧评分的影响。通过在引入局部人工失真的实验中测试多组帧级质量预测器的定位性能,结果表明其表现优于众包感知实验中的人类标注检测能力。
原文摘要 · Abstract (English)
While automatic subjective speech quality assessment has witnessed much progress, an open question is whether an automatic quality assessment at frame resolution is possible. This would be highly desirable, as it adds explainability to the assessment of speech synthesis systems. Here, we take first steps towards this goal by identifying issues of existing quality predictors that prevent sensible frame-level prediction. Further, we define criteria that a frame-level predictor should fulfill. We also suggest a chunk-based processing that avoids the impact of a localized distortion on the score of neighboring frames. Finally, we measure in experiments with localized artificial distortions the localization performance of a set of frame-level quality predictors and show that they can outperform detection performance of human annotations obtained from a crowd-sourced perception experiment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。