arXiv:2608.09571cs.SD2026-08

用分块路由专家模型生成统一音频场景,让语音音乐声效协同可控。

SonicWeave: Chunk-Routed Mixture-of-Experts for Unified Audio Scene Generation

论文配图:SonicWeave: Chunk-Routed Mixture-of-Experts for Unified Audio Scene Generation
图 1 · 摘自论文原文
  • 按音频块分组路由专家,结合全局文本提示与局部声学状态决策。
  • 在多类音频合成任务中优于基线模型,复杂场景生成更连贯。
  • 适合需要混合音源统一生成的研究者或音频系统开发者。

文本驱动的通用音频生成正从孤立的语音、音乐和音效合成,转向单一模型生成可控且连贯的音频场景。这一统一设定极具挑战:异质成分对共享骨干网络提出矛盾的结构要求,而复杂混合场景可能包含局部差异或重叠内容,需在同一片段内实现细粒度适应。现有音频混合专家(MoE)主要在领域层面路由,而逐标记路由忽略了声学信号固有的连续性。本文提出SonicWeave,一种用于统一音频场景生成的流匹配模型。核心是分块路由的混合专家(CPE-MoE),通过冲突门控先验-证据路由机制实现。该机制结合全局先验(编码结构化文本条件与扩散阶段信息)与局部演化声学状态的证据,当局部状态不可靠时优先采用先验,当区域偏离全局场景上下文时则允许局部证据影响路由。SonicWeave支持语音、音乐、音效、歌唱及其细粒度混合,仅使用一组权重。在TTS、TTA和TTM基准测试中,SonicWeave持续优于受控密集模型和基础MoE基线。复杂场景评估显示生成组合质量提升,路由分析揭示扩散阶段中专家根据内容呈现专业化。结果表明,时间连贯的先验-证据路由是统一音频生成的有效条件计算策略。

原文摘要 · Abstract (English)

Text-conditioned general audio generation is moving beyond isolated speech, music, and sound-effect synthesis toward a single model that can compose them into controllable, coherent audio scenes. This unified setting is particularly challenging: heterogeneous components impose conflicting structural requirements on a shared backbone, while a complex mixed scene may contain locally distinct or overlapping content that demands fine-grained adaptation within the same clip. Existing audio mixture-of-experts (MoEs) mainly route at the domain level, while token-wise routing overlooks the local continuity inherent to acoustic signals. We propose SonicWeave, a flow-matching model for unified audio scene generation. At its core is a chunk-routed MoE with a conflict-gated prior-evidence routing mechanism (CPE-MoE). CPE-MoE routes contiguous acoustic chunks by combining a global prior that encodes the structured text condition and diffusion phase with local evidence from the evolving acoustic state. A learned conflict gate favors the prior when local states are unreliable, while allowing local evidence to influence routing when a region departs from the global scene context. SonicWeave supports speech, music, sound effects, singing, and their fine-grained mixtures with a single set of weights. Across TTS, TTA, and TTM benchmarks, SonicWeave consistently improves over controlled Dense and Base-MoE baselines. Complex-scene evaluation further demonstrates improved compositional quality, while routing analyses reveal content-dependent expert specialization across diffusion phases. These results suggest that temporally coherent, prior-evidence routing is an effective conditional-computation strategy for unified audio generation. Project page: https://caiyunrui.github.io/SonicWeave.

音频生成混合专家流匹配统一建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。