用分块路由专家模型生成统一音频场景,让语音音乐声效协同可控。
SonicWeave: Chunk-Routed Mixture-of-Experts for Unified Audio Scene Generation

- 按音频块分组路由专家,结合全局文本提示与局部声学状态决策。
- 在多类音频合成任务中优于基线模型,复杂场景生成更连贯。
- 适合需要混合音源统一生成的研究者或音频系统开发者。
文本驱动的通用音频生成正从孤立的语音、音乐和音效合成,转向单一模型生成可控且连贯的音频场景。这一统一设定极具挑战:异质成分对共享骨干网络提出矛盾的结构要求,而复杂混合场景可能包含局部差异或重叠内容,需在同一片段内实现细粒度适应。现有音频混合专家(MoE)主要在领域层面路由,而逐标记路由忽略了声学信号固有的连续性。本文提出SonicWeave,一种用于统一音频场景生成的流匹配模型。核心是分块路由的混合专家(CPE-MoE),通过冲突门控先验-证据路由机制实现。该机制结合全局先验(编码结构化文本条件与扩散阶段信息)与局部演化声学状态的证据,当局部状态不可靠时优先采用先验,当区域偏离全局场景上下文时则允许局部证据影响路由。SonicWeave支持语音、音乐、音效、歌唱及其细粒度混合,仅使用一组权重。在TTS、TTA和TTM基准测试中,SonicWeave持续优于受控密集模型和基础MoE基线。复杂场景评估显示生成组合质量提升,路由分析揭示扩散阶段中专家根据内容呈现专业化。结果表明,时间连贯的先验-证据路由是统一音频生成的有效条件计算策略。
原文摘要 · Abstract (English)
Text-conditioned general audio generation is moving beyond isolated speech, music, and sound-effect synthesis toward a single model that can compose them into controllable, coherent audio scenes. This unified setting is particularly challenging: heterogeneous components impose conflicting structural requirements on a shared backbone, while a complex mixed scene may contain locally distinct or overlapping content that demands fine-grained adaptation within the same clip. Existing audio mixture-of-experts (MoEs) mainly route at the domain level, while token-wise routing overlooks the local continuity inherent to acoustic signals. We propose SonicWeave, a flow-matching model for unified audio scene generation. At its core is a chunk-routed MoE with a conflict-gated prior-evidence routing mechanism (CPE-MoE). CPE-MoE routes contiguous acoustic chunks by combining a global prior that encodes the structured text condition and diffusion phase with local evidence from the evolving acoustic state. A learned conflict gate favors the prior when local states are unreliable, while allowing local evidence to influence routing when a region departs from the global scene context. SonicWeave supports speech, music, sound effects, singing, and their fine-grained mixtures with a single set of weights. Across TTS, TTA, and TTM benchmarks, SonicWeave consistently improves over controlled Dense and Base-MoE baselines. Complex-scene evaluation further demonstrates improved compositional quality, while routing analyses reveal content-dependent expert specialization across diffusion phases. These results suggest that temporally coherent, prior-evidence routing is an effective conditional-computation strategy for unified audio generation. Project page: https://caiyunrui.github.io/SonicWeave.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。