为普什图语语音合成设计自动化评估框架,解决低资源非拉丁文字语言评估难题。
PashtoTTS-Bench: automated screening for low-resource non-Latin-script text-to-speech

- 提出INVS-A框架,从可听性、自然度、文本保真度和语言识别四方面自动评估
- 在200个FLEURS和200个筛选过的Common Voice数据上测试,OmniVoice auto WER最低(FLEURS 24.1%)
- 适合低资源语言语音合成研究者,提供完整评估工具与失败日志
低资源非拉丁字母语言的语音合成(TTS)评估常因单一自动语音识别(ASR)回环词错误率(WER)而失效。系统可能无音频输出、说出邻近语言、仅在ASR转录中保留目标文字,或对母语者听起来不自然。本文提出INSV(可理解性、自然度、文本保真度、验证)报告框架,将这些情况分离。本论文发布INSV-A:自动化筛查子集,包括合成完成度、ASR WER/CER、转录文本保真率及音频语言识别。母语者主观评分与音位标注虽已定义但未在此版本中提供。将INSV-A实例化为PashtoTTS-Bench,用于2026年4-5月的普什图语TTS评测。测试了Edge GulNawaz、Edge Latifa、OmniVoice clone、OmniVoice auto及乌尔都语负向对照,使用200个FLEURS和200个过滤后的Common Voice 24提示。在独立omniASR_CTC_300M_v2下,OmniVoice auto的WER最低(FLEURS 24.1%,CV24 27.4%),其次为Edge GulNawaz(32.8%,39.5%)、Edge Latifa(35.6%,47.7%)、OmniVoice clone(45.4%,34.8%)。WER低于自然语音基线反映合成音频清晰,不代表优于真人发音。Whisper Large V3在检测的普什图语合成音频上返回0.0%普什图标签,而MMS-LID-4017与SpeechBrain VoxLingua107能将普什图输出与乌尔都语对照区分开。本次发布包含提供商元数据、每句得分、语言识别审计、失败日志及系统添加脚本。
原文摘要 · Abstract (English)
Text-to-speech (TTS) evaluation for low-resource non-Latin-script languages can fail when it relies on a single ASR round-trip word error rate (WER). A system may produce no audio, speak a neighbouring language, preserve target script text only in an ASR transcript, or sound unnatural to native listeners. We introduce INSV (Intelligibility, Naturalness, Script fidelity, and Verification), a reporting framework that separates these cases. This paper reports INSV-A, the automated screening subset: synthesis completion, ASR WER/CER, transcript Script Fidelity Rate, and audio language identification. Native MOS and phonetic annotation are specified but not claimed in this release. We instantiate INSV-A as PashtoTTS-Bench, a dated benchmark for Pashto TTS. The April-May 2026 run evaluates Edge GulNawaz, Edge Latifa, OmniVoice clone, OmniVoice auto, and an Urdu negative control on 200 FLEURS and 200 filtered Common Voice 24 prompts. Under the independent omniASR_CTC_300M_v2, OmniVoice auto has the lowest WER (24.1% FLEURS, 27.4% CV24), followed by Edge GulNawaz (32.8%, 39.5%), Edge Latifa (35.6%, 47.7%), and OmniVoice clone (45.4%, 34.8%). WER below the natural-speech baseline reflects clean synthetic audio and should not be read as better than native speech. Whisper Large V3 returns 0.0% Pashto labels on checked Pashto TTS audio, while MMS-LID-4017 and SpeechBrain VoxLingua107 separate Pashto outputs from the Urdu control. The release provides provider metadata, per-sentence scores, LID audits, failure logs, and scripts for adding systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。