arXiv:2505.23781cs.SDcs.LG2025-05被引 2

融合降噪与深度特征,实现高精度语音异常检测

Unified AI for Accurate Audio Anomaly Detection

  • 用谱减法和自适应滤波提升音频质量
  • 结合MFCC与OpenL3嵌入,提升特征表达能力
  • 多模型集成在TORGO和LibriSpeech上表现优异

本文提出一种统一的AI框架,通过整合先进的降噪、特征提取与机器学习建模技术,实现高精度音频异常检测。方法首先采用谱减法与自适应滤波提升音频质量,随后利用传统方法(如MFCC)及预训练模型(如OpenL3)提取深度特征。建模阶段融合经典模型(SVM、随机森林)、深度网络(CNN)与集成方法,增强鲁棒性与准确率。在TORGO和LibriSpeech等基准数据集上评估,该框架在精确率、召回率以及口齿不清与正常语音分类上均表现优异,有效应对嘈杂环境与实时应用挑战,提供可扩展的音频异常检测方案。

原文摘要 · Abstract (English)

This paper presents a unified AI framework for high-accuracy audio anomaly detection by integrating advanced noise reduction, feature extraction, and machine learning modeling techniques. The approach combines spectral subtraction and adaptive filtering to enhance audio quality, followed by feature extraction using traditional methods like MFCCs and deep embeddings from pre-trained models such as OpenL3. The modeling pipeline incorporates classical models (SVM, Random Forest), deep learning architectures (CNNs), and ensemble methods to boost robustness and accuracy. Evaluated on benchmark datasets including TORGO and LibriSpeech, the proposed framework demonstrates superior performance in precision, recall, and classification of slurred vs. normal speech. This work addresses challenges in noisy environments and real-time applications and provides a scalable solution for audio-based anomaly detection.

音频异常检测深度特征多模型集成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。