轻量实时音乐分离模型,延迟更低、参数更少,适合听障设备等场景。
Towards Practical Real-Time Low-Latency Music Source Separation
- 采用单路径结构与通道扩展融合,提升实时处理效率。
- 参数量显著减少,推理时间比现有模型快,满足低延迟要求。
- 适用于助听器、现场演出等对延迟敏感的应用场景。
近年来,深度学习在音乐分离领域取得显著进展。然而,针对实时、低延迟音乐分离的研究仍较少,这类技术在助听器、音频流重混音和现场演出中具有潜在应用价值。此外,大型模型的持续发展限制了其在部分场景的应用。本文提出一种轻量级实时低延迟模型RT-STT,基于双路径TFC-TDF UNET(DTTNet)改进而来。通过引入基于通道扩展的特征融合方法,并验证单路径建模在实时模型中的优势。同时研究量化方法以进一步降低推理时间。实验表明,相比当前最优模型,RT-STT在参数量更少、推理时间更短的情况下,性能表现更优。
原文摘要 · Abstract (English)
In recent years, significant progress has been made in the field of deep learning for music demixing. However, there has been limited attention on real-time, low-latency music demixing, which holds potential for various applications, such as hearing aids, audio stream remixing, and live performances. Additionally, a notable tendency has emerged towards the development of larger models, limiting their applicability in certain scenarios. In this paper, we introduce a lightweight real-time low-latency model called Real-Time Single-Path TFC-TDF UNET (RT-STT), which is based on the Dual-Path TFC-TDF UNET (DTTNet). In RT-STT, we propose a feature fusion technique based on channel expansion. We also demonstrate the superiority of single-path modeling over dual-path modeling in real-time models. Moreover, we investigate the method of quantization to further reduce inference time. RT-STT exhibits superior performance with significantly fewer parameters and shorter inference times compared to state-of-the-art models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。