统一生成长时序音频,支持多角色、多场景混合音频合成。
Qwen-Audio-3.0-Gen-Preview Technical Report

- 采用扩散Transformer与共享VAE,非自回归生成完整混音波形。
- 在多说话人与复杂时间线任务中表现优异,跨轮一致性与定位准确。
- 适合需要统一处理多种音频任务的开发者和研究者。
现有单领域与多任务音频系统在直接组织异构音频成分、环境音效及多个角色构建长时序场景方面仍存在局限。我们提出 Qwen-Audio-3.0-Gen-Preview,一个统一的非自回归框架,利用扩散Transformer(DiT)与共享变分自编码器(VAE)生成完整的混合波形。提示增强将自由文本请求转化为结构化时间记录,并作为文本条件进行渲染;两阶段数据课程与语义条件视图训练模型在独立与混合场景音频中使用这些条件。共享连续VAE将48kHz立体声波形压缩为25Hz隐变量序列,并引入语义监督,实现异构音频的统一表示。在公开参考条件基准上,该模型在所有三个子集中的说话人相似性表现最为突出;在多说话人与丰富时间线基准上,其显著优势分别体现在双语言跨轮一致性和时间定位能力;在AudioCaps上,其优势集中于使用大音频-语言模型与AudioBox的评估结果。这些结果证明了无需任务专用分支即可实现时序结构化音频统一生成的潜力。
原文摘要 · Abstract (English)
Existing single-domain and multi-task audio systems remain limited in directly organizing heterogeneous audio components, ambience, and multiple roles into long-form temporal scenes. We present Qwen-Audio-3.0-Gen-Preview, a unified non-autoregressive framework that uses a Diffusion Transformer (DiT) and a shared variational autoencoder (VAE) to generate the complete mixed waveform. Prompt enhancement converts free-form requests into structured temporal records that are rendered as textual conditions, while a two-stage data curriculum and semantic conditional views train the proposed model to use these conditions across standalone and mixed-scene audio. A shared continuous VAE compresses 48kHz stereo waveforms into 25Hz latent sequences and incorporates semantic supervision, providing one representation for heterogeneous audio. On the public reference-conditioned benchmark, speaker similarity is the proposed model's clearest strength across all three subsets. Across the multi-speaker and rich-timeline benchmarks, its clearest comparative strengths are cross-turn consistency in both languages and temporal localization, respectively. On AudioCaps, its advantages are concentrated in evaluations using large audio-language models and AudioBox. These results demonstrate the potential of unified generation for temporally structured audio without task-specific branches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。