对比不同位置编码对时频域双路径模型分离效果与长序列泛化能力的影响
A Comparative Study on Positional Encoding for Time-frequency Domain Dual-path Transformer-based Source Separation Models
- 在训练长度内使用位置编码提升分离性能
- 无位置编码模型在长序列外推上表现更优,尤其含卷积层时
- 为长音频分离任务选择编码方式提供实证依据
本研究探讨了位置编码(PE)对基于时频域双路径Transformer模型的语音源分离性能及长序列外推能力的影响。由于该类模型在处理长时音频输入和未见采样率信号时,其外推能力至关重要,而位置编码对此影响显著,但现有研究对此关注不足。为此,我们以最新的先进模型TF-Locoformer为基础,对比多种位置编码方法。结果表明:(i) 在与训练序列长度相同或更短的情况下,带位置编码的模型性能更优;(ii) 但不使用位置编码的模型展现出更强的长度外推能力,这一趋势在包含卷积层的模型中尤为明显。
原文摘要 · Abstract (English)
In this study, we investigate the impact of positional encoding (PE) on source separation performance and the generalization ability to long sequences (length extrapolation) in Transformer-based time-frequency (TF) domain dual-path models. The length extrapolation capability in TF-domain dual-path models is a crucial factor, as it affects not only their performance on long-duration inputs but also their generalizability to signals with unseen sampling rates. While PE is known to significantly impact length extrapolation, there has been limited research that explores the choice of PEs for TF-domain dual-path models from this perspective. To address this gap, we compare various PE methods using a recent state-of-the-art model, TF-Locoformer, as the base architecture. Our analysis yields the following key findings: (i) When handling sequences that are the same length as or shorter than those seen during training, models with PEs achieve better performance. (ii) However, models without PE exhibit superior length extrapolation. This trend is particularly pronounced when the model contains convolutional layers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。