arXiv:2604.22209eess.AScs.AI2026-04ACL被引 4

一个能用文字指令统一生成语音、音乐和音效的模型。

UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions

论文配图:UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions
图 1 · 摘自论文原文
  • 用动态标记注入将杂乱声音转为结构化时序表示,支持精确时长控制。
  • 在语音合成与音乐生成上分别达到WER 1.47%和SongEval Coherence 3.18的领先水平。
  • 多任务联合训练带来正向迁移,提升结构连贯性与语调表现力。

生成式音频建模长期被分割为语音合成(TTS)、音乐生成(TTM)和音效生成(TTA)等专用任务,各自采用不同的控制范式。由于结构性语义表征(语音/音乐)与非结构化声学纹理(音效)之间存在本质差异,实现跨模态统一仍是核心挑战。本文提出UniSonate,一种统一的流匹配框架,可通过标准化、无参考的自然语言指令生成语音、音乐与音效。为解决结构差异,我们设计动态标记注入机制,将非结构化环境音映射至结构化时序潜在空间,从而在以音素驱动的多模态扩散变换器(MM-DiT)中实现精确时长控制。结合多阶段课程学习策略,有效缓解跨模态优化冲突。大量实验表明,UniSonate在基于指令的TTS(WER 1.47%)与TTM(SongEval Coherence 3.18)上达到当前最优性能,同时在TTA中保持竞争力。关键发现:在多样化音频数据上的联合训练显著增强结构连贯性与语调表现力,优于单任务基线。音频样例见 https://qiangchunyu.github.io/UniSonate/。

原文摘要 · Abstract (English)

Generative audio modeling has largely been fragmented into specialized tasks, text-to-speech (TTS), text-to-music (TTM), and text-to-audio (TTA), each operating under heterogeneous control paradigms. Unifying these modalities remains a fundamental challenge due to the intrinsic dissonance between structured semantic representations (speech/music) and unstructured acoustic textures (sound effects). In this paper, we introduce UniSonate, a unified flow-matching framework capable of synthesizing speech, music, and sound effects through a standardized, reference-free natural language instruction interface. To reconcile structural disparities, we propose a novel dynamic token injection mechanism that projects unstructured environmental sounds into a structured temporal latent space, enabling precise duration control within a phoneme-driven Multimodal Diffusion Transformer (MM-DiT). Coupled with a multi-stage curriculum learning strategy, this approach effectively mitigates cross-modal optimization conflicts. Extensive experiments demonstrate that UniSonate achieves state-of-the-art performance in instruction-based TTS (WER 1.47%) and TTM (SongEval Coherence 3.18), while maintaining competitive fidelity in TTA. Crucially, we observe positive transfer, where joint training on diverse audio data significantly enhances structural coherence and prosodic expressiveness compared to single-task baselines. Audio samples are available at https://qiangchunyu.github.io/UniSonate/.

语音生成音乐生成统一模型文本控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。