arXiv:2508.06405eess.ASeess.SP2025-08中稿 · ICASSP 2026

用硬标签法让模型自动判断语音是否平稳,速度快精度高。

Acoustic Non-Stationarity Objective Assessment with Hard Label Criteria for Supervised Learning Models

  • 设计硬标签准则生成全局平稳性标签,替代传统耗时评估方法。
  • 提出NANSA模型,分类准确率达99%,显著优于现有方法。
  • 适合需要实时语音分析的场景,如语音识别与通信系统。

客观非平稳性度量方法资源消耗大,难以用于实时处理。本文提出一种新型硬标签准则(HLC)算法,为声学信号生成全局非平稳性标签,使监督学习模型可被训练为平稳性估计器。首先在主流通用声学模型上验证了其有效性,表明这些模型已具备捕捉平稳性信息的能力。进一步提出首个基于HLC的声学非平稳性评估网络(NANSA),该模型性能超越现有方法,在测试中达到最高99%的分类准确率,同时解决了传统客观度量方法计算不可行的问题。

原文摘要 · Abstract (English)

Objective non-stationarity measures are resource intensive and impose critical limitations for real-time processing solutions. In this paper, a novel Hard Label Criteria (HLC) algorithm is proposed to generate a global non-stationarity label for acoustic signals, enabling supervised learning strategies to be trained as stationarity estimators. The HLC is first evaluated on state-of-the-art general-purpose acoustic models, demonstrating that these models capture stationarity information. Furthermore, the first-of-its-kind HLC-based Network for Acoustic Non-Stationarity Assessment (NANSA) is proposed. NANSA models outperform competing approaches, achieving up to 99% classification accuracy, while solving the computational infeasibility of traditional objective measures.

语音分析监督学习非平稳性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。