arXiv:2605.31530eess.AScs.SD2026-05被引 2

一个模型搞定语音生成、声音合成与音频编辑,效率更高。

UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion

论文配图:UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion
图 1 · 摘自论文原文
  • 用多层大模型融合注入语义信息,提升指令理解能力。
  • 621万至732万参数,性能超越专用模型且体积缩小4倍。
  • 适合需要统一处理多种音频任务的研究者与开发者。

我们提出UNISON,一种统一的声学生成与编辑框架,基于潜在扩散模型,将语音生成、声音生成与音频编辑整合于单一模型中。该模型可处理文本到音频、文本到语音、零样本说话人克隆、混合语音与声音生成、场景级音频编辑、语音在场景中的编辑及定时时间组合等任务,所有任务共享一组权重。其核心设计包括:(1) 层级深度大模型融合,通过学习投影将冻结多模态大模型(MLLM)均匀采样层的隐藏状态注入对应多模态迪特(MM-DiT)模块,实现深度匹配的语义条件控制,显著优于单层基线;(2) 统一多任务架构,仅通过通道掩码编码任务身份,并以VAE编码通道拼接提供源音频。训练通过在线GPU端多任务数据合成流水线、任务同质化批处理和两阶段课程学习实现稳定。模型参数量为621M–732M,跨领域表现媲美甚至超过专用模型,同时约为同类统一系统的1/4大小。

原文摘要 · Abstract (English)

We present UNISON, a latent diffusion framework that unifies speech generation, sound generation, and audio editing within a single model. A single model handles text-to-audio, text-to-speech, zero-shot speaker cloning, mixed speech-and-sound generation, scene-level audio editing, speech-in-scene editing, and timed temporal composition, all of which share a single set of weights. Our architecture features two core designs: (1) Layer-wise deep LLM fusion, which injects hidden states from uniformly sampled layers of a frozen MLLM into corresponding MM-DiT blocks via learned projections, providing depth-matched semantic conditioning that improves instruction following over single-layer baselines; and (2) a unified multi-task architecture where task identity is encoded solely by a channel-wise mask and source audio is provided through VAE-encoded channel concatenation. Training is stabilized by an online GPU-side multi-task data synthesis pipeline with task-homogeneous batching and a two-stage curriculum. With 621M--732M trainable parameters, UNISON achieves results competitive with or exceeding task-specialist models across evaluated domains, while being roughly $4\times$ smaller than comparable unified systems.

音频生成大模型融合统一框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。