arXiv:2410.08235cs.SDcs.LG2024-10被引 1

用迁移学习提升语音通话中机器答录机识别准确率

A Recurrent Neural Network Approach to the Answering Machine Detection Problem

  • 基于YAMNet提取特征,用循环网络实现实时音频流分类
  • 测试集准确率达96%以上,结合静音检测可超98%
  • 适合语音客服、外呼系统等需要实时判别的场景

在电信与云通信领域,实时准确判断外拨电话是否由真人或答录机接听至关重要。该问题在营销推广中尤为关键,能提升服务品质、效率并降低成本。尽管意义重大,现有研究仍不充分。本文提出一种创新方法,利用YAMNet模型进行特征提取,并构建基于循环神经网络的分类器,实现对音频流的实时处理,而非固定长度录音。实验结果表明,测试集准确率超过96%。进一步分析误分类样本发现,结合FFmpeg提供的静音检测算法,准确率可提升至98%以上。

原文摘要 · Abstract (English)

In the field of telecommunications and cloud communications, accurately and in real-time detecting whether a human or an answering machine has answered an outbound call is of paramount importance. This problem is of particular significance during campaigns as it enhances service quality, efficiency and cost reduction through precise caller identification. Despite the significance of the field, it remains inadequately explored in the existing literature. This paper presents an innovative approach to answering machine detection that leverages transfer learning through the YAMNet model for feature extraction. The YAMNet architecture facilitates the training of a recurrent-based classifier, enabling real-time processing of audio streams, as opposed to fixed-length recordings. The results demonstrate an accuracy of over 96% on the test set. Furthermore, we conduct an in-depth analysis of misclassified samples and reveal that an accuracy exceeding 98% can be achieved with the integration of a silence detection algorithm, such as the one provided by FFmpeg.

语音识别机器学习通信系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。