用图分割方法精确定位口吃起始点,提升语音治疗反馈精度
StutterCut: Uncertainty-Guided Normalised Cut for Dysfluency Segmentation
- 将语音片段建模为图节点,通过不确定性引导的图划分实现精准分割
- 在真实与合成数据上均达更高F1值,口吃起始点检测更准确
- 扩展了FluencyBank数据集,提供四类口吃帧级标注,更具现实意义
口吃检测与分割对有效语音治疗和实时反馈至关重要。然而,现有方法多仅在语句层面进行分类。本文提出StutterCut,一种半监督框架,将口吃分割建模为图划分问题:来自重叠语音窗口的嵌入表示为图节点。利用基于弱标签(语句级)训练的伪真值分类器优化节点连接,其影响由蒙特卡洛丢弃带来的不确定性度量控制。同时,我们扩展了弱标签数据集FluencyBank,新增四种口吃类型的帧级边界标注,构建更贴近现实的基准。在真实与合成数据上的实验表明,StutterCut优于现有方法,在F1分数和口吃起始点检测精度上均有提升。
原文摘要 · Abstract (English)
Detecting and segmenting dysfluencies is crucial for effective speech therapy and real-time feedback. However, most methods only classify dysfluencies at the utterance level. We introduce StutterCut, a semi-supervised framework that formulates dysfluency segmentation as a graph partitioning problem, where speech embeddings from overlapping windows are represented as graph nodes. We refine the connections between nodes using a pseudo-oracle classifier trained on weak (utterance-level) labels, with its influence controlled by an uncertainty measure from Monte Carlo dropout. Additionally, we extend the weakly labelled FluencyBank dataset by incorporating frame-level dysfluency boundaries for four dysfluency types. This provides a more realistic benchmark compared to synthetic datasets. Experiments on real and synthetic datasets show that StutterCut outperforms existing methods, achieving higher F1 scores and more precise stuttering onset detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。