arXiv:2608.12715cs.SDcs.AI2026-08

混合域声学增强模型,用动态路由提升降噪效果与推理效率

HybridSB-MoE: Dual-Domain Schrödinger Bridges with Scene-Adaptive Expert Routing for Speech Enhancement

论文配图:HybridSB-MoE: Dual-Domain Schrödinger Bridges with Scene-Adaptive Expert Routing for Speech Enhancement
图 1 · 摘自论文原文
  • 双域融合:频谱路径捕捉不确定性,波形桥模型处理随机噪声
  • 专家路由机制:5种不同架构的MoE,top-k=2实现场景自适应选择
  • 理论保障:推导出有限步采样误差上界,小步数推理有数学保证

生成式语音增强存在三大缺陷:频谱模型虽能捕捉谐波结构但常破坏相位,波形模型保留相位却丢失谐波,而薛定谔桥(SB)虽缩短了从噪声到清晰语音的变换路径,但推理成本与训练关联松散。我们提出 HybridSB-MoE,一个统一的双域框架,通过三项贡献解决上述问题。首先,非对称不确定性融合:频谱路径通过专家分歧捕获认知不确定性,波形桥通过随机动力学建模偶然性变异,二者非对称融合,使混合权重可自适应不同错误场景而非简单平均。其次,异构MoE设计采用 top-k=2 路由策略,覆盖五种不同架构原型,架构多样性使认知信号反映归纳偏置失效而非微小扰动。第三,离散化误差界(定理1):路径一致性与轨迹正则项共同将 K 步桥采样误差在 2-沃瑟斯坦距离下界为 K^{-alpha},使得小步数推理具备可证明的性能保证,而非仅经验性宣称。在 VoiceBank+DEMAND 数据集上,HybridSB-MoE 在相同步数预算下优于基于扩散和 SB 的基线方法,同时保持与一致性蒸馏的少步方法相当的竞争力。

原文摘要 · Abstract (English)

Generative speech enhancement faces three gaps: spectral models capture harmonic structure but often disrupt phase, waveform models preserve phase but miss harmonics, and Schrödinger Bridges (SB) shorten transport from noise to clean speech but leave inference cost only loosely tied to training. We propose HybridSB-MoE, a dual-domain framework that fills these gaps through three contributions unified by a single asymmetric design principle. (i) Asymmetric uncertainty fusion: The spectral path captures epistemic uncertainty via expert disagreement, while the waveform bridge models aleatoric variance through stochastic dynamics. We fuse them asymmetrically, allowing the mixing weight to adapt to distinct error regimes rather than average predictions. (ii) Heterogeneous MoE with top-k=2 routing across five distinct architectural archetypes, where architectural diversity makes the epistemic signal indicate which inductive bias fails rather than small perturbations among similar experts. (iii) Discretization bound (Theorem 1): path-consistency and trajectory regularizers together bound the K-step bridge sampling error in 2-Wasserstein distance at rate K-alpha, making small-K inference an objective-level guarantee rather than an empirical claim. On VoiceBank+DEMAND, HybridSB-MoE outperforms diffusion- and SB-based baselines at their step budgets while remaining competitive with consistency-distilled few-step methods.

语音增强MoESchrödinger桥生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。