arXiv:2410.04797cs.SDcs.MM2024-10

通过注意力融合多层级语音特征,提升语音障碍诊断准确率。

Attentive-based Multi-level Feature Fusion for Voice Disorder Diagnosis

  • 用ECAPA-TDNN和Wav2vec 2.0提取语音深层病理特征
  • 注意力机制引导双模型特征在多层融合,提升诊断性能
  • 在FEMH和SVD数据集上分别达到90.51%和87.68%准确率

语音障碍严重影响日常生活质量。然而,从原始音频中准确识别病理类别仍面临挑战,主要因数据集有限。本文提出一种新框架,通过在潜在空间融合多层次病理信息,实现全面的语音病理特征提取。模型采用两阶段训练:首先使用在多个领域表现优异的ECAPA-TDNN和Wav2vec 2.0从原始音频中学习通用病理特征;其次设计注意力融合模块,分别对ECAPA-TDNN与Wav2vec 2.0提取的病理特征进行交互建模并指导多层融合,整个模型基于预训练特征联合微调,以完成自动语音病理检测任务。在FEMH和SVD数据集上的大量实验表明,该框架优于现有基线方法,在两个数据集上分别达到90.51%和87.68%的准确率。

原文摘要 · Abstract (English)

Voice disorders negatively impact the quality of daily life in various ways. However, accurately recognizing the category of pathological features from raw audio remains a considerable challenge due to the limited dataset. A promising method to handle this issue is extracting multi-level pathological information from speech in a comprehensive manner by fusing features in the latent space. In this paper, a novel framework is designed to explore the way of high-quality feature fusion for effective and generalized detection performance. Specifically, the proposed model follows a two-stage training paradigm: (1) ECAPA-TDNN and Wav2vec 2.0 which have shown remarkable effectiveness in various domains are employed to learn the universal pathological information from raw audio; (2) An attentive fusion module is dedicatedly designed to establish the interaction between pathological features projected by EcapTdnn and Wav2vec 2.0 respectively and guide the multi-layer fusion, the entire model is jointly fine-tuned from pre-trained features by the automatic voice pathology detection task. Finally, comprehensive experiments on the FEMH and SVD datasets demonstrate that the proposed framework outperforms the competitive baselines, and achieves the accuracy of 90.51% and 87.68%.

语音诊断特征融合注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。