通过教师-学生微调,有效抑制神经网络吉他失真模型的频谱混叠。
Anti-aliasing of neural distortion effects via model fine tuning
- 用教师-学生框架微调模型,学习无混叠的输出特征。
- 在多数场景下,去混叠效果优于两倍过采样,尤其对LSTM和TCN有效。
- 适合追求高保真音色还原的音乐信号处理研究者使用。
近年来,神经网络在吉他失真效果建模中广泛应用。尽管其能生成听觉上逼真的模型,但在高频高增益输入下易产生频率混叠。非线性激活函数在扩展信号带宽时,不仅生成期望的谐波失真,也引入了非期望的混叠失真。本文提出一种基于教师-学生微调的方法,其中教师为权重冻结的预训练模型,学生为其可学习参数的复制品。学生模型在由正弦信号通过原模型并移除非谐波成分生成的无混叠数据集上进行微调。实验结果表明,该方法显著抑制了长短期记忆网络(LSTM)和时间卷积网络(TCN)中的混叠现象。在多数案例中,其去混叠效果优于两倍过采样。该方法的一个副作用是谐波失真成分也被影响,但此影响具有模型依赖性,其中LSTM模型在去混叠与保持模拟设备听感相似性之间取得了最佳平衡。
原文摘要 · Abstract (English)
Neural networks have become ubiquitous with guitar distortion effects modelling in recent years. Despite their ability to yield perceptually convincing models, they are susceptible to frequency aliasing when driven by high frequency and high gain inputs. Nonlinear activation functions create both the desired harmonic distortion and unwanted aliasing distortion as the bandwidth of the signal is expanded beyond the Nyquist frequency. Here, we present a method for reducing aliasing in neural models via a teacher-student fine tuning approach, where the teacher is a pre-trained model with its weights frozen, and the student is a copy of this with learnable parameters. The student is fine-tuned against an aliasing-free dataset generated by passing sinusoids through the original model and removing non-harmonic components from the output spectra. Our results show that this method significantly suppresses aliasing for both long-short-term-memory networks (LSTM) and temporal convolutional networks (TCN). In the majority of our case studies, the reduction in aliasing was greater than that achieved by two times oversampling. One side-effect of the proposed method is that harmonic distortion components are also affected. This adverse effect was found to be model-dependent, with the LSTM models giving the best balance between anti-aliasing and preserving the perceived similarity to an analog reference device.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。