优化音频分离模型,速度提升81%且性能损失极小。
FasTUSS: Faster Task-Aware Unified Source Separation
- 改进双路径时频模型结构,减少计算量
- 速度提升81%,性能仅下降1.2dB
- 适合需要高效实时音频处理的场景
当前表现最佳的音频源分离网络架构为时频(TF)双路径模型,在语音增强、音乐分离和电影音频分离任务中达到顶尖水平。尽管参数量较低,但仍需大量运算,导致执行时间较长。随着大模型趋势发展,如最近提出的任务感知统一源分离(TUSS)模型,该问题更加突出。TUSS基于TF-Locoformer构建,通过一系列提示词指定需分离的音源数量与类型,实现单一条件模型解决多种任务。本文分析了TUSS的设计选择,旨在优化性能与复杂度的权衡。提出两个更高效的模型FasTUSS-8.3G和FasTUSS-11.7G,分别将原始模型运算量降低81%和73%,在所有基准测试上平均性能仅下降1.2~dB和0.4~dB。此外,研究提示词条件的影响,推导出因果版本的TUSS模型。
原文摘要 · Abstract (English)
Time-Frequency (TF) dual-path models are currently among the best performing audio source separation network architectures, achieving state-of-the-art performance in speech enhancement, music source separation, and cinematic audio source separation. While they are characterized by a relatively low parameter count, they still require a considerable number of operations, implying a higher execution time. This problem is exacerbated by the trend towards bigger models trained on large amounts of data to solve more general tasks, such as the recently introduced task-aware unified source separation (TUSS) model. TUSS, which aims to solve audio source separation tasks using a single, conditional model, is built upon TF-Locoformer, a TF dual-path model combining convolution and attention layers. The task definition comes in the form of a sequence of prompts that specify the number and type of sources to be extracted. In this paper, we analyze the design choices of TUSS with the goal of optimizing its performance-complexity trade-off. We derive two more efficient models, FasTUSS-8.3G and FasTUSS-11.7G that reduce the original model's operations by 81\% and 73\% with minor performance drops of 1.2~dB and 0.4~dB averaged over all benchmarks, respectively. Additionally, we investigate the impact of prompt conditioning to derive a causal TUSS model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。