arXiv:2604.13288cs.CLcs.AI2026-04

用双语宪法文本训练,让濒危的克丘亚语也能自然朗读。

Giving Voice to the Constitution: Low-Resource Text-to-Speech for Quechua and Spanish Using a Bilingual Legal Corpus

  • 用跨语言迁移,用西班牙语数据帮克丘亚语补足语音数据
  • 三种先进语音合成模型均实现高质量双语输出
  • 成果开源,适合做原住民语言技术研究者参考

我们提出一个统一的语音合成流程,基于三种前沿文本转语音(TTS)架构(XTTS v2、F5-TTS 和 DiFlow-TTS),为秘鲁宪法内容生成高质量的克丘亚语和西班牙语语音。模型分别在异质规模与录音条件的西班牙语和克丘亚语语音数据集上独立训练,利用双语与多语言TTS能力提升两种语言的合成质量。通过跨语言迁移,该框架缓解了克丘亚语数据稀缺问题,同时保持西班牙语语音的自然度。我们公开了每个宪法条文的训练检查点、推理代码和合成音频,为原住民语言及多语言环境中的语音技术提供可复用资源。本工作推动了低资源场景下政治与法律内容的包容性语音系统发展。

原文摘要 · Abstract (English)

We present a unified pipeline for synthesizing high-quality Quechua and Spanish speech for the Peruvian Constitution using three state-of-the-art text-to-speech (TTS) architectures: XTTS v2, F5-TTS, and DiFlow-TTS. Our models are trained on independent Spanish and Quechua speech datasets with heterogeneous sizes and recording conditions, and leverage bilingual and multilingual TTS capabilities to improve synthesis quality in both languages. By exploiting cross-lingual transfer, our framework mitigates data scarcity in Quechua while preserving naturalness in Spanish. We release trained checkpoints, inference code, and synthesized audio for each constitutional article, providing a reusable resource for speech technologies in indigenous and multilingual contexts. This work contributes to the development of inclusive TTS systems for political and legal content in low-resource settings.

语音合成原住民语言低资源双语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。