用片段标签训练模型,实现精准的口吃片段定位与分类
Stuttering Classification and Segmentation with Attention-Based Multiple Instance Learning

- 基于注意力机制的多实例学习,从片段级标签推断帧级信息
- 帧级F1提升23%,片段级F1提升2%-9%
- 适合语音病理分析、言语康复研究者使用
利用深度学习进行口吃检测与分类具有提升口吃严重程度评估效率的潜力。现有大多数口吃分类数据集仅提供片段级标签,难以满足确定单个口吃不流畅持续时间所需的细粒度帧级分类需求。为此,我们提出一种基于微调wav2vec 2.0、WavLM和Whisper编码器的多实例神经网络架构,采用实例级与嵌入级多实例学习方法,在片段级标注数据上训练模型,完成片段级与帧级口吃分类任务。实验结果表明,帧级F1分数提升23%,片段级F1分数提升2%至9%,证明了模型能够有效利用片段级数据实现帧级分割。
原文摘要 · Abstract (English)
Stuttering detection and classification using deep learning methods has the potential to improve the process of stuttering severity assessment. Most stuttering classification datasets provide clip-level labels, making them unsuitable for fine-grained frame-level classification needed to determine the duration of individual stuttering dysfluencies. To overcome this challenge, we present a multiple instance neural network architecture based on fine-tuned wav2vec 2.0, WavLM and Whisper encoders. We apply instance- and embedding-based multiple instance learning approaches to train models on a clip-level dataset for both clip-level and frame-level stuttering classification tasks. Our results show a 23% improvement in frame-level F1 score and between 2% and 9% in clip-level F1 score, demonstrating the ability of our models to utilize clip-level data for frame-level segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。