arXiv:2605.28063cs.SDcs.AI2026-05

用自然语言直接生成语音与音效融合的统一音频

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts

论文配图:Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts
图 1 · 摘自论文原文
  • 基于大模型自回归框架,无需外部文本改写
  • 在语音、音效及混合场景中均优于现有方法
  • 适合需要灵活创作复合音频的创作者

音频生成取得显著进展,但如何从自由文本直接合成包含语音与音效及其自然融合的统一音频仍是挑战。现有方法或依赖分离流程,无法捕捉细粒度交互;或需结构化输入与外部文本重写,限制了自由文本提示的灵活性。本文提出新任务:自由文本提示到统一音频生成,旨在直接从自然语言合成包含语音、音效及其复合内容的统一音频。为此,我们提出PlanAudio,一种基于大语言模型的统一自回归框架。首先,利用大模型内在推理能力简化架构,替代传统文本编码器;其次,引入语义潜空间思维链机制,作为隐式规划桥梁,连接高层语义理解与底层声学合成。此外,我们构建PlanAudio-Bench,一个专门评估复合音频场景的基准。在语音、音效及其复合场景下进行评估,结果表明PlanAudio普遍优于现有流程与统一基线,同时在单场景模型中保持竞争力。分析进一步揭示语义潜空间思维链优于其他思维链机制,并强调连续多场景训练课程的重要性。

原文摘要 · Abstract (English)

Audio generation has made significant progress, yet synthesizing unified audio where speech and sounds are naturally composited remains a challenge. Current methods either rely on disjoint pipelines, which fail to capture fine-grained interactions, or require structured inputs and external text rewriting, which limits the flexibility of free-form text prompts. In this paper, we introduce a new task: Free-Form-Text-Prompt-to-Unified-Audio generation, which aims to directly synthesize unified audio containing speech, sound, and their composites from unconstrained natural language. To address this task, we propose PlanAudio, a unified, autoregressive LLM-based framework. First, it simplifies the model architecture by leveraging intrinsic LLM reasoning capability instead of traditional text encoders. Second, it introduces a semantic latent chain-of-thought mechanism, an implicit planning mechanism that bridges high-level semantic understanding and low-level acoustic synthesis. Furthermore, we create PlanAudio-Bench, a specialized benchmark for evaluating composite audio scenarios. We perform evaluations in the scenarios of speech, sound, and their composites. The results demonstrate that PlanAudio generally outperforms the existing pipeline and unified baselines, while staying competitive with models designed for a single scenario. Our analysis further reveals the superiority of semantic latent CoT over other CoT mechanisms and highlights the importance of continuous multi-scenario training curricula.

语音生成音效合成大模型统一音频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。