2bit量化会严重损害语音验证性能,该研究提出新方法有效缓解误差。
On Low-Bit Quantization Errors in Speaker Verification: Diagnostic and Mitigation

- 通过分层与得分级分析,定位量化误差源头
- 2比特时出现明显性能拐点,决策错误集中于浮点阈值附近
- 多精度级联方案仅对模糊样本升级,兼顾效率与精度
尽管低比特量化为在资源受限设备上部署语音验证提供了可行方案,但其对性能的影响仍不明确。本文通过联合分层与得分级分析,研究了基于均匀K均值的ResNet-36和ResNet-200的量化感知训练。分层分析揭示了脆弱组件,表明得分退化不能仅由权重失真解释;在2比特处出现明显拐点,得分漂移与有害决策翻转集中在FP32阈值附近。得分级分析进一步揭示极端量化下得分误差的产生位置与机制。基于此,提出校准的多精度级联策略,大部分样本在2比特下完成推理,仅对模糊案例升级,实现接近FP32的性能,同时显著降低计算与内存开销。
原文摘要 · Abstract (English)
Although low-bit quantization provides practical means to deploy speaker verification on resource-constrained devices, its effects on speaker verification performance remain poorly understood. In this paper, we study uniform K-means quantization-aware training of ResNet-36 and ResNet-200 through joint layer-wise and score-level analyses. Our layer-wise analysis highlights fragile components and shows that score degradation is not fully explained by weight distortion alone. We identify a clear knee point at 2 bits, with larger score drift and harmful decision flips concentrated near the FP32 threshold. Our score-level analysis reveals where and how score errors emerge under extreme quantization. Building on these findings, we propose a calibrated multi-precision cascade that resolves most trials at 2 bits and escalates only ambiguous cases, achieving performance close to FP32 while preserving the efficiency benefits of low-bit inference with substantially lower compute and memory costs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。