提出高效评估神经编解码器量化效果的新方法,节省训练成本。
Efficient Evaluation of Quantization-Effects in Neural Codecs
- 用低复杂度模型模拟大网络的非线性行为,实现快速评估。
- 发现梯度传递技术对系统性能有显著影响,验证了不同方法差异。
- 适合研究量化效应或优化训练效率的工程师与研究员。
神经编解码器(包含编码器、量化器和解码器)可在极低比特率下实现信号传输。训练此类系统需使用直通估计器、软到硬退火或统计量化器模拟等技术,以在量化器处维持非零梯度。评估量化影响(如梯度传递技术对整体系统的影响)通常代价高昂且耗时,因需大量训练且缺乏低成本可靠的评估指标。本文提出一种基于模拟数据的高效评估框架,采用指定比特数的低复杂度神经编码器/解码器,模拟大网络中的非线性行为。该系统在训练时间、计算量和硬件需求上均高度高效,使我们能够揭示神经编解码器的多种行为特征。基于发现,我们改进了直通估计器的训练稳定性。结果在内部神经音频编解码器及当前最先进描述性音频编解码器(descript-audio-codec)上得到验证。
原文摘要 · Abstract (English)
Neural codecs, comprising an encoder, quantizer, and decoder, enable signal transmission at exceptionally low bitrates. Training these systems requires techniques like the straight-through estimator, soft-to-hard annealing, or statistical quantizer emulation to allow a non-zero gradient across the quantizer. Evaluating the effect of quantization in neural codecs, like the influence of gradient passing techniques on the whole system, is often costly and time-consuming due to training demands and the lack of affordable and reliable metrics. This paper proposes an efficient evaluation framework for neural codecs using simulated data with a defined number of bits and low-complexity neural encoders/decoders to emulate the non-linear behavior in larger networks. Our system is highly efficient in terms of training time and computational and hardware requirements, allowing us to uncover distinct behaviors in neural codecs. We propose a modification to stabilize training with the straight-through estimator based on our findings. We validate our findings against an internal neural audio codec and against the state-of-the-art descript-audio-codec.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。