ASR回环评估会掩盖中文新闻语音中因语境导致的朗读错误,需结合人工审计。
ASR-Roundtrip Evaluation Can Mask Context- and Convention-Dependent Reading Errors in Chinese News TTS
- 通过隔离语段诊断,发现46例被ASR误判为正确的错误朗读
- 在110个高风险案例中,46例被ASR掩盖、51例无错误
- 适合关注中文语音合成质量评估的研究者与评测人员
ASR回环评估被广泛用作文本转语音(TTS)可懂性的可扩展代理,但可能产生听者感知到的朗读错误的假阴性。本文研究了中文新闻语音中依赖上下文或领域惯例正确朗读的片段,如体育比分、飞机型号、技术单位和会员名称。在这些情况下,原始TTS可能选择看似合理但错误的读法,而ASR却将其转录为预期或表面正确的文本。针对110个高风险MiMo TTS案例的定向审计(含完整基数)确认:46例被掩盖的假阴性、9例暴露的TTS错误、55例无错误。语段隔离诊断重新暴露了其中18/46例先前被掩盖的错误。对同一数据集进行仅基于Raw TTS的CosyVoice审计,确认51例被掩盖的错误。在97个经双重审计确认被掩盖的TTS音频中,Qwen3-ASR表面恢复了40例,而Paraformer仅恢复2例。结果表明,ASR回环评估适用于筛查,但不足以作为中文新闻朗读风险评估的独立真实标准。
原文摘要 · Abstract (English)
ASR-roundtrip evaluation is widely used as a scalable proxy for text-to-speech (TTS) intelligibility, but it can produce false negatives for reading errors perceived by listeners. We study Chinese news TTS spans whose correct reading depends on context or domain conventions, such as sports scores, aircraft models, technical units, and membership names. In these cases, Raw TTS can choose a plausible but wrong reading while ASR transcribes the audio as the intended or surface-correct text. A targeted audit over 110 high-risk MiMo TTS cases, reported with a complete denominator, confirms 46 masked false negatives, 9 exposed TTS errors, and 55 cases with no Raw TTS error. A span-isolation diagnostic re-exposes 18/46 previously masked errors. A Raw-only CosyVoice audit on the same targeted pool confirms 51 masked cases. Across the 97 TTS-specific audio files labeled confirmed masked across the two audits, Qwen3-ASR surface-recovers 40 cases, whereas Paraformer does so in only 2. The results suggest that ASR-roundtrip is useful for screening but insufficient as standalone ground truth for Chinese news reading-risk evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。