统一音频生成框架,用离散码本令牌实现多任务高质量音频合成
UniTok-Audio: A Unified Audio Generation Framework via Generative Modeling on Discrete Codec Tokens
- 基于自回归方式将条件特征转为音频离散令牌
- 在五项时序对齐任务中表现媲美顶尖专用模型
- 双流编解码器提升波形重建保真度,适合多任务研究者
生成建模在文本、图像和音频领域均取得显著进展,展现出强大的统一表征学习能力。然而,当前音频生成模型在音质和跨任务泛化能力方面仍面临挑战,导致开发重复、性能不一且难以扩展。为此,我们提出 extbf{UniTok-Audio},一个可扩展的统一音频生成框架。具体包括:1)将条件连续特征转化为目标音频的离散令牌,采用自回归方式生成;2)通过特殊任务标识符令牌,在单一框架中统一多种任务的学习模式;3)设计包含声学与语义分支的双流音频编解码器,实现高保真波形重建。实验表明,UniTok-Audio 在五项时序对齐任务(语音修复、目标说话人提取、语音分离、语音转换、语言查询音频源分离)上表现优于或媲美现有最先进单任务或多任务系统。为推动后续研究,我们将开源代码库。演示页面见:https://alibaba.github.io/unified-audio。
原文摘要 · Abstract (English)
Generative modeling has recently achieved remarkable success across text, image, and audio domains, demonstrating powerful capabilities for unified representation learning. However, audio generation models still face challenges in terms of audio quality and generalization ability across tasks. This fragmentation results in redundant development efforts, inconsistent performance, and limited extensibility. To address these issues, we propose \textbf{UniTok-Audio}, a scalable and extensible framework for unified audio generation tasks. Specifically, 1) UniTok-Audio extracts continuous feature of conditions to generates discrete tokens of target audio in an autoregressive manner; 2) a special task identifier token unifies different learning patterns of multiple tasks in a single framework; 3) a dual-stream audio codec involving acoustic and semantic branch is developed for high-fidelity waveform reconstruction. Experimental results demonstrate that UniTok-Audio achieves competitive performance in comparation with state-of-the-art task-specific or multi-task systems across five time-aligned tasks: speech restoration, target speaker extraction, speech separation, voice conversion, and language-queried audio source separation. To foster future research, we will open-source our codebase. The demo page of our work can be found here: https://alibaba.github.io/unified-audio.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。