arXiv:2606.29575cs.SDcs.AI2026-06中稿 · Interspeech 2026

TF-MoE通过时频动态专家选择,实现低算力下的语音分离性能提升。

TF-MoE: Time-Frequency Mixture-of-Experts for Efficient Speech Separation

论文配图:TF-MoE: Time-Frequency Mixture-of-Experts for Efficient Speech Separation
图 1 · 摘自论文原文
  • 在时频维度交替使用动态专家选择,实现稀疏激活
  • 在4.1 GMACs/s算力下,比BSRNN高3.8 dB SDR
  • 适合资源受限的边缘设备语音分离场景

近年来语音分离(SS)模型虽趋于轻量化,但推理计算成本仍限制其在边缘设备上的部署。为此,我们提出TF-MoE——一种稀疏混合专家(MoE)框架,可在几乎不增加推理开销的前提下提升模型容量。该方法通过交替使用时域和频域的MoE模块,在每帧或梅尔频带级别动态选择专家,实现时频维度上的动态专家专精。基于梅尔频带分割的Conformer骨干网络,TF-MoE在低算力条件下表现出色。实验表明,在保持约4.1 GMACs/s推理成本的前提下,其在Libri2Mix数据集上相较BSRNN实现+3.8 dB的SDR提升,验证了其在边缘设备部署中的潜力。

原文摘要 · Abstract (English)

Recent advances in speech separation (SS) have led to compact front-end models with small parameter sizes, yet their high computational cost remains a major barrier for deployment on edge devices. To address this, we propose TF-MoE, a sparse Mixture-of-Experts (MoE) framework that enhances model capacity with almost no increase in inference cost. Our method introduces dynamic expert specialization in time and frequency dimensions through alternating time-wise and frequency-wise MoE modules, each dynamically selecting experts per frame or mel band. Built upon a mel-band-splitting Conformer backbone, TF-MoE achieves strong performance on SS tasks under low-compute settings. Experimental results demonstrate that TF-MoE consistently improves separation performance under computation cost constraints, outperforming BSRNN by +3.8 dB SDR on Libri2Mix with comparable 4.1 GMACs/s inference cost. This positions TF-MoE as a promising candidate for edge-device deployment.

语音分离MoE边缘计算高效模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。