arXiv:2603.28086cs.SDcs.AI2026-03被引 4

用自然语言生成真实感语音,让声音更贴近真人表现。

MOSS-VoiceGenerator: Create Realistic Voices with Natural Language Descriptions

  • 基于自然语言指令直接生成语音音色,无需录音
  • 在影视级语料上训练,提升声音真实感与表现力
  • 适合角色配音、虚拟助手等需要个性化语音的场景

从自然语言生成语音旨在直接根据自由文本描述生成说话人音色,使用户能为特定角色、个性或情绪定制声音。这种可控制的声音生成对故事讲述、游戏配音、角色扮演代理和对话助手等下游应用具有重要意义,是现代文本转语音模型的重要任务。然而,现有模型主要在精心录制的录音棚数据上训练,生成的语音虽清晰准确,却缺乏真实人类声音的生活质感。为此,我们提出MOSS-VoiceGenerator,一个开源的指令驱动语音生成模型,可直接从自然语言提示生成新音色。受‘接触真实声学变化能产生更自然感知声音’的假设启发,我们在大规模来自影视内容的表达性语音数据上进行训练。主观偏好测试表明,该模型在整体性能、指令遵循能力及自然度方面均优于其他语音设计模型。

原文摘要 · Abstract (English)

Voice design from natural language aims to generate speaker timbres directly from free-form textual descriptions, allowing users to create voices tailored to specific roles, personalities, and emotions. Such controllable voice creation benefits a wide range of downstream applications-including storytelling, game dubbing, role-play agents, and conversational assistants, making it a significant task for modern Text-to-Speech models. However, existing models are largely trained on carefully recorded studio data, which produces speech that is clean and well-articulated, yet lacks the lived-in qualities of real human voices. To address these limitations, we present MOSS-VoiceGenerator, an open-source instruction-driven voice generation model that creates new timbres directly from natural language prompts. Motivated by the hypothesis that exposure to real-world acoustic variation produces more perceptually natural voices, we train on large-scale expressive speech data sourced from cinematic content. Subjective preference studies demonstrate its superiority in overall performance, instruction-following, and naturalness compared to other voice design models.

语音生成自然语言个性化声音AI配音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。