arXiv:2601.21886eess.AS2026-01中稿 · ICASSP 2026被引 3

通过一致性约束提升语音质量评分的可解释性

Speech Quality-Based Localization of Low-Quality Speech and Text-to-Speech Synthesis Artefacts

  • 用分段一致性约束优化整体评分模型,降低逐帧波动
  • 在两个主流语音合成系统中成功识别低质语音片段
  • 结果经听觉测试验证,低帧评分段更易被判定为劣质

大量研究从话语或系统层面评估语音质量,虽能判断整体优劣,却难以解释评分依据。帧级评分更具可解释性,但因训练时缺乏强监督信号,模型难以调优和正则化。本文提出利用分段一致性约束对话语级语音质量预测模型进行正则化,显著降低帧级评分的随机性。进一步展示了两项应用:部分伪造场景下的定位,以及在两种先进文本到语音合成系统中检测合成伪影。通过听觉测试发现,由低帧评分定义的片段被听众评定为低质量的比例,远高于随机对照组。

原文摘要 · Abstract (English)

A large number of works view the automatic assessment of speech from an utterance- or system-level perspective. While such approaches are good in judging overall quality, they cannot adequately explain why a certain score was assigned to an utterance. frame-level scores can provide better interpretability, but models predicting them are harder to tune and regularize since no strong targets are available during training. In this work, we show that utterance-level speech quality predictors can be regularized with a segment-based consistency constraint which notably reduces frame-level stochasticity. We then demonstrate two applications involving frame-level scores: The partial spoof scenario and the detection of synthesis artefacts in two state-of-the-art text-to-speech systems. For the latter, we perform listening tests and confirm that listeners rate segments to be of poor quality more often in the set defined by low frame-level scores than in a random control set.

语音质量可解释性语音合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。