arXiv:2512.07393cs.LG2025-12

调优截断反向传播提升音频效果模型训练效率与质量

Empirical Results for Adjusting Truncated Backpropagation Through Time while Training Neural Audio Effects

  • 通过调整序列数、批大小和序列长度优化TBPTT训练
  • 模型准确率提升,计算开销降低,训练更稳定
  • 适合音频处理、神经网络建模方向的研究者

本文研究了在数字音频效果建模中使用截断反向传播通过时间(TBPTT)训练神经网络的优化方法,重点关注动态范围压缩。通过卷积-循环架构,在有无用户控制条件的数据集上进行了大量实验,评估了序列数、批大小和序列长度等关键超参数的影响。结果表明,合理调参可提升模型准确性与训练稳定性,同时降低计算需求。客观评估显示优化设置下性能改善,主观听感测试也证实新配置保持高感知质量。

原文摘要 · Abstract (English)

This paper investigates the optimization of Truncated Backpropagation Through Time (TBPTT) for training neural networks in digital audio effect modeling, with a focus on dynamic range compression. The study evaluates key TBPTT hyperparameters -- sequence number, batch size, and sequence length -- and their influence on model performance. Using a convolutional-recurrent architecture, we conduct extensive experiments across datasets with and without conditionning by user controls. Results demonstrate that carefully tuning these parameters enhances model accuracy and training stability, while also reducing computational demands. Objective evaluations confirm improved performance with optimized settings, while subjective listening tests indicate that the revised TBPTT configuration maintains high perceptual quality.

音频建模TBPTT神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。