arXiv:2606.24714cs.CLcs.SD2026-06

评测中文新闻语音合成对复杂文本的发音准确性

CN-NewsTTS Bench: a target-level automatic benchmark for raw-input Chinese news TTS pronunciation

  • 基于原始文本自动评估语音合成对数字、缩写等目标的发音
  • 最佳系统准确率达0.879,部分系统低于0.60
  • 适合关注中文语音合成真实场景表现的研究者

中文新闻文本包含大量密集书面形式,如数字、连字符命名、范围、单位符号、百分比、英文缩写及中英混合名称。这些形式在实际听觉流程中频繁出现,而语音合成系统在保留原文字符串的同时可能改变发音含义。我们提出 CN-NewsTTS Bench v0.1,一个开放的目标级基准,用于评估中文新闻语音合成产品能否从原始文本中正确发音,无需用户规则、大模型重写、SSML提示或手动编辑。发布包含200条开发集、800条公开测试集、992个可自动评估的目标,固定转录来自三引擎语音识别集成、自动评分器,以及七个产品系统的初始结果。此外报告了ASR路径诊断、子集消融实验、类别级结果、置信区间及厂商配置元数据。最佳系统严格准确率为0.879,多个系统仍低于0.60。

原文摘要 · Abstract (English)

Chinese news text contains dense written forms such as scores, hyphenated model names, ranges, unit symbols, percentages, English abbreviations, and mixed Chinese-Latin-digit names. These forms are frequent in real listening workflows, and a text-to-speech (TTS) system can preserve the written string while changing the spoken meaning. We introduce CN-NewsTTS Bench v0.1, an open target-level benchmark for evaluating whether Chinese news TTS products pronounce such targets correctly from raw text, without user-side rules, LLM rewriting, SSML hints, or manual edits. The release contains a 200-record development set, an 800-record public test set, 992 public auto-evaluable targets, fixed transcripts from a three-ASR ensemble, an automatic target scorer, and initial results for seven product TTS systems. We additionally report ASR-route diagnostics, ASR-subset ablations, category-level results, confidence intervals, and provider configuration metadata. The best system reaches 0.879 strict accuracy, while several systems remain below 0.60.

语音合成中文语音自动评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。