神经音频编码器在低带宽下表现优于传统编码,但高比特率时会损失说话人辨识特征。
Analysis of Speaker Verification Performance Trade-offs with Neural Audio Codec Transmission
- 对比传统与神经音频编码器在不同码率下的说话人验证性能
- 低码率(<12kbps)下神经编码器比Opus提升6-8%,高码率(≈24kbps)时误识率仅上升0.4-0.7%
- 适合关注低带宽语音通信的系统设计者,提示需开发面向说话人的编码器
近年来,神经音频编码器(NACs)在音频处理中广泛应用,但可能引入失真,影响说话人验证(SV)性能。本研究在VoxCeleb1数据集上评估了多种先进SV模型在不同码率下使用传统和神经音频编码器的表现。结果表明,所有模型和编码器在码率下降时均出现性能持续下降。值得注意的是,与传统编码器相比,神经编码器并未根本破坏SV性能;在低码率(<12 kbps)下,其性能比Opus高出6-8%;在较高码率(≈24 kbps)下略逊于Opus,EER仅增加0.4-0.7%。这种差距可能源于NACs主要优化听觉感知质量,无意中丢弃了关键的说话人区分特征,而Opus则专门保留语音特性。研究建议,NACs在带宽受限场景下是可行替代方案,未来应发展说话人感知的编码器或对SV模型进行再训练适配。
原文摘要 · Abstract (English)
Neural audio codecs (NACs) have made significant advancements in recent years and are rapidly being adopted in many audio processing pipelines. However, they can introduce audio distortions which degrade speaker verification (SV) performance. This study investigates the impact of both traditional and neural audio codecs at varying bitrates on three state of-the-art SV models evaluated on the VoxCeleb1 dataset. Our findings reveal a consistent degradation in SV performance across all models and codecs as bitrates decrease. Notably, NACs do not fundamentally break SV performance when compared to traditional codecs. They outperform Opus by 6-8% at low-bitrates (< 12 kbps) and remain marginally behind at higher bitrates ($\approx$ 24 kbps), with an EER increase of only 0.4-0.7%. The disparity at higher bitrates is likely due to the primary optimization of NACs for perceptual quality, which can inadvertently discard critical speaker-discriminative features, unlike Opus which was designed to preserve vocal characteristics. Our investigation suggests that NACs are a feasible alternative to traditional codecs, especially under bandwidth limitations. To bridge the gap at higher bitrates, future work should focus on developing speaker-aware NACs or retraining and adapting SV models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。