arXiv:2606.20137eess.AScs.CL2026-06中稿 · INTERSPEECH 2026

专注音调重音错误检测,提升语音质量评估精度

PASQA: Pitch-Accent-Focused Speech Quality Assessment Model Trained on Synthetic Speech with Accent Errors

论文配图:PASQA: Pitch-Accent-Focused Speech Quality Assessment Model Trained on Synthetic Speech with Accent Errors
图 1 · 摘自论文原文
  • 基于自监督表示与音节条件融合,显式建模音调重音错误
  • 在有无重音错误的语音上均实现高排序准确率,优于传统模型
  • 适合需要精准评估日语发音质量的研究与应用

现有话语级自然度评分模型通常对局部音调重音错误不敏感。本文提出面向音调重音的语音质量评估模型PASQA,专门关注音调重音正确性。通过使用可控重音的文本转语音系统构建受控的日语重音错误数据集,并根据重音错误率计算伪重音质量分数。PASQA基于自监督表征,采用音节条件融合、排序损失、辅助重音错误定位任务和说话人无关训练。实验表明,传统模型无法保持重音错误严重程度的排序,而PASQA在已见和未见说话人上均取得高排序准确性,且与人工重音正确性判断具有更强一致性。代码已开源。

原文摘要 · Abstract (English)

Existing mean opinion score (MOS) prediction models typically predict utterance-level naturalness MOS and can be insensitive to localized pitch-accent errors. We propose Pitch-Accent-focused Speech Quality Assessment (PASQA), which explicitly targets pitch-accent correctness. To train our model, we construct a controlled Japanese accent-error dataset by changing accent patterns using an accent-controllable text-to-speech system, and compute a pseudo accent-quality score from the accent-error rate. PASQA builds on self-supervised representations and employs mora-conditioned fusion, ranking loss, an auxiliary accent-error localization task, and speaker-invariant training. Experiments show that conventional models fail to preserve the ordering by accent-error severity, whereas PASQA achieves high ordering accuracy on both seen and unseen speakers. Further, PASQA shows stronger agreement with human accent-correctness judgments. The code is available at https://github.com/lycorp-jp/PASQA.

语音评估音调重音自监督日语语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。