首个濒危女书语音合成系统,用五级声调标记提升发音还原度。
NüshuVoice: Reviving the Voice of Endangered Nüshu with Pitch-Aware Text-to-Speech

- 基于声调标注设计条件化语音合成框架,利用声调作为韵律先验。
- 在极低资源下实现高保真音色与准确声调重建,人类评测可懂度达92%。
- 适合语言保护、小语种语音技术研究者使用,开源数据集已发布。
女书是一种中国湖南江永县女性历史上使用的濒危拼音文字。现有计算研究多集中于文本数字化与视觉识别,其真实发音的声学重建仍鲜有探索。构建女书文本转语音(TTS)系统面临极大挑战:可用录音极少,且多为孤立音节而非自然语句。本文提出首个女书语音合成基准NüshuVoice,构建了包含标准化Unicode女书文本、音标转写、标准汉语翻译及档案录音的句子级女书语音数据集。针对极端低资源场景,提出Nüshu-PitchVITS,一种以基频(F0)为条件的VITS框架,显式利用女书五级声调标记作为韵律先验。实验表明,该方法在频谱保真度、声调还原和人工评测可懂度上均优于强基线模型。数据集与代码已公开:https://anonymous.4open.science/r/Nvshu-TTS-2EB6。
原文摘要 · Abstract (English)
Nüshu is an endangered phonetic script historically used by women in Jiangyong County, southern Hunan, China. While existing computational studies of Nüshu mainly focus on textual digitization and visual recognition, the acoustic reconstruction of its authentic pronunciation remains largely unexplored. Building a Nüshu text-to-speech (TTS) system is particularly challenging because available recordings are extremely limited and mostly consist of isolated syllable-level pronunciations rather than natural sentence-level utterances. In this work, we introduce NüshuVoice, the first TTS benchmark for Nüshu. We construct a sentence-level Nüshu text-to-audio dataset that aligns standardized Unicode Nüshu text, phonetic transcriptions, standard Chinese translations, and archival recordings. To synthesize speech under this extreme low-resource setting, we propose Nüshu-PitchVITS, an F0-conditioned VITS framework that leverages Nüshu's five-level pitch notation as an explicit prosodic inductive bias. Experimental results show that Nüshu-PitchVITS outperforms strong TTS baselines in spectral fidelity, pitch reconstruction, and human-rated intelligibility. We publicly release the dataset and code at: https://anonymous.4open.science/r/Nvshu-TTS-2EB6.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。