arXiv:2502.02703cs.CLcs.AI2025-02NAACL被引 5

用轻量流匹配模型实现三种原住民语言的多语言语音合成

Developing multilingual speech synthesis system for Ojibwe, Mi'kmaq, and Maliseet

  • 采用无注意力机制的轻量流匹配架构
  • 三语联合训练在数据少时优于单语模型
  • 强调社区参与的人类评估重要性

我们为北美三种原住民语言——奥吉布瓦语、米克马克语和马利塞特语,开发了轻量级流匹配多语言文本到语音系统。结果表明,在数据稀缺情况下,对三种类型相似的语言进行联合训练,可显著提升语音合成性能,优于单语模型。无注意力架构在性能上与自注意力架构相当,且内存效率更高。本研究不仅推动了低资源语言复兴的技术进展,也揭示了当前人类评估协议中存在的文化偏差,呼吁采用更以社区为中心的评估方法。

原文摘要 · Abstract (English)

We present lightweight flow matching multilingual text-to-speech (TTS) systems for Ojibwe, Mi'kmaq, and Maliseet, three Indigenous languages in North America. Our results show that training a multilingual TTS model on three typologically similar languages can improve the performance over monolingual models, especially when data are scarce. Attention-free architectures are highly competitive with self-attention architecture with higher memory efficiency. Our research not only advances technical development for the revitalization of low-resource languages but also highlights the cultural gap in human evaluation protocols, calling for a more community-centered approach to human evaluation.

语音合成低资源语言原住民语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。