arXiv:2604.27607cs.CL2026-04被引 1

JaiTTS实现高保真泰语语音克隆,支持混合语言输入且表现超越真人

JaiTTS: A Thai Voice Cloning Model

论文配图:JaiTTS: A Thai Voice Cloning Model
图 1 · 摘自论文原文
  • 基于VoxCPM架构,无需文本归一化直接处理数字与泰英混用
  • 短时语音合成字错误率1.94%,优于真人基准1.98%;长时任务持平真人
  • 人评胜出283/400次,显著领先商业旗舰模型,适合本地化语音应用

我们提出JaiTTS-v1.0,一种基于大型泰语语料库持续训练的先进泰语语音克隆文语转换模型。该模型架构源自无分词器的自回归式TTS模型VoxCPM,可直接处理数字及泰英混用文本,无需显式文本归一化。我们在短时与长时语音生成任务上测试模型,覆盖多种真实应用场景。JaiTTS-v1.0在短时任务中实现1.94%的字错误率(CER),优于人类参考值1.98%;长时任务表现与人类参考值相当。在人工评估中,模型在400组两两对比中赢得283次,仅输58次。代码与演示已开源于https://github.com/JTS-AI-Team/JaiTTS。

原文摘要 · Abstract (English)

We present JaiTTS-v1.0, a state-of-the-art Thai voice cloning text-to-speech model built through continual training on a large Thai-centric speech corpus. The model architecture is adapted from VoxCPM, a tokenizer-free autoregressive TTS model. JaiTTS-v1.0 directly processes numerals and Thai-English code-switching, which is very common in realistic settings, without explicit text normalization. We test the models on short- and long-duration speech generation, which reflects many real-world use cases. JaiTTS-v1.0 achieves a state-of-the-art CER of 1.94%, surpassing the human ground truth of 1.98% for short-duration tasks while performing on par with human ground truth for long-duration tasks. In human judgment evaluations, our model wins 283 of 400 pairwise comparisons against commercial flagships, with only 58 losses. Our code and demo are available at https://github.com/JTS-AI-Team/JaiTTS .

语音克隆泰语生成自回归模型多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。