arXiv:2410.21478cs.SDcs.AI2024-10被引 1

用轻量模型加速语音通话早期数据分类,兼顾速度与准确率。

Knowledge Distillation for Real-Time Classification of Early Media in Voice Communications

  • 用梯度提升树替代卷积网络,降低资源消耗。
  • 推理速度显著提升,准确率与复杂模型相当。
  • 适合部署在资源受限的实时语音分析场景。

本文研究语音通话初始化阶段早期媒体的实时分类工业场景。现有方法多依赖卷积神经网络,但存在资源开销大的问题。本文提出基于梯度提升树的轻量化方案,结合知识蒸馏与类别聚合技术,训练更小、更快的模型。实验在私有和公开数据集上验证,新方法在保持可比准确率的同时显著提升运行效率。此外,在印度某区域数据中心的实际案例中,性能提升明显,证明其在真实生产环境中的可行性。

原文摘要 · Abstract (English)

This paper investigates the industrial setting of real-time classification of early media exchanged during the initialization phase of voice calls. We explore the application of state-of-the-art audio tagging models and highlight some limitations when applied to the classification of early media. While most existing approaches leverage convolutional neural networks, we propose a novel approach for low-resource requirements based on gradient-boosted trees. Our approach not only demonstrates a substantial improvement in runtime performance, but also exhibits a comparable accuracy. We show that leveraging knowledge distillation and class aggregation techniques to train a simpler and smaller model accelerates the classification of early media in voice calls. We provide a detailed analysis of the results on a proprietary and publicly available dataset, regarding accuracy and runtime performance. We additionally report a case study of the achieved performance improvements at a regional data center in India.

语音分类轻量化模型知识蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。