arXiv:2509.15666cs.SDcs.AI2025-09

通过动态推理重复实现音源分离的高效可扩展框架

TISDiSS: A Training-Time and Inference-Time Scalable Framework for Discriminative Source Separation

  • 采用早期分裂多损失监督与共享参数设计,支持训练与推理时扩展
  • 减少参数量同时达成领先性能,低延迟场景下效果更优
  • 适合需要灵活速度-精度权衡的实时音源分离应用

音源分离是语音、音乐和音频处理中的基础任务,也为生成模型训练提供更清晰的数据。然而,提升分离性能通常依赖更大网络,导致训练和部署成本上升。受生成建模中推理时扩展的启发,我们提出训练与推理均可扩展的判别式音源分离框架TISDiSS,融合早期分裂多损失监督、共享参数设计与动态推理重复机制。TISDiSS通过调整推理深度即可实现灵活的速度-性能权衡,无需重新训练。系统分析表明,增加推理重复次数能提升浅层推理表现,有利于低延迟应用。在标准语音分离基准上,TISDiSS以更少参数量实现最优性能,证明其可扩展性与实用性。代码已开源:https://github.com/WingSingFung/TISDiSS。

原文摘要 · Abstract (English)

Source separation is a fundamental task in speech, music, and audio processing, and it also provides cleaner and larger data for training generative models. However, improving separation performance in practice often depends on increasingly large networks, inflating training and deployment costs. Motivated by recent advances in inference-time scaling for generative modeling, we propose Training-Time and Inference-Time Scalable Discriminative Source Separation (TISDiSS), a unified framework that integrates early-split multi-loss supervision, shared-parameter design, and dynamic inference repetitions. TISDiSS enables flexible speed-performance trade-offs by adjusting inference depth without retraining additional models. We further provide systematic analyses of architectural and training choices and show that training with more inference repetitions improves shallow-inference performance, benefiting low-latency applications. Experiments on standard speech separation benchmarks demonstrate state-of-the-art performance with a reduced parameter count, establishing TISDiSS as a scalable and practical framework for adaptive source separation. Code is available at https://github.com/WingSingFung/TISDiSS.

音源分离可扩展性动态推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。