arXiv:2508.08715eess.AScs.AI2025-08被引 6

MultiGen用大模型生成适合儿童的多语言语音,支持低资源语种。

MultiGen: Child-Friendly Multilingual Speech Generator with LLMs

  • 基于大模型架构实现适龄多语言语音生成
  • 在三种低资源语言中表现优于基线方法
  • 适合儿童语言学习与跨文化互动场景

生成式语音模型在提升人机交互方面展现出巨大潜力,尤其在儿童语言学习等实际应用中价值显著。然而,针对低资源语言在多样文化背景下的高质量、儿童友好型语音生成仍具挑战。本文提出MultiGen,一种基于大模型架构的多语言儿童友好语音生成模型,专为低资源语言设计。该模型通过适龄多语言语音生成,支持新加坡口音普通话、马来语和泰米尔语,让儿童能以符合文化背景的方式与AI交流。客观指标与主观评估结果均表明,MultiGen在性能上显著优于基线方法。

原文摘要 · Abstract (English)

Generative speech models have demonstrated significant potential in improving human-machine interactions, offering valuable real-world applications such as language learning for children. However, achieving high-quality, child-friendly speech generation remains challenging, particularly for low-resource languages across diverse languages and cultural contexts. In this paper, we propose MultiGen, a multilingual speech generation model with child-friendly interaction, leveraging LLM architecture for speech generation tailored for low-resource languages. We propose to integrate age-appropriate multilingual speech generation using LLM architectures, which can be used to facilitate young children's communication with AI systems through culturally relevant context in three low-resource languages: Singaporean accent Mandarin, Malay, and Tamil. Experimental results from both objective metrics and subjective evaluations demonstrate the superior performance of the proposed MultiGen compared to baseline methods.

语音生成多语言儿童友好大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。