arXiv:2606.08210eess.AScs.CL2026-06中稿 · INTERSPEECH 2026

用分层图神经网络捕捉儿童发音特征,提升发育性口吃识别能力

Paediatric-HGNN: A Hybrid Heterogeneous Graph Neural Network for Detecting Disfluency in Children's Speech via Multiscale Acoustic Fusion

论文配图:Paediatric-HGNN: A Hybrid Heterogeneous Graph Neural Network for Detecting Disfluency in Children's Speech via Multiscale Acoustic Fusion
图 1 · 摘自论文原文
  • 构建词与声学片段的异构图,建模层级交互关系
  • 在两个儿科语料库上达到82.4%加权准确率,典型口吃F1达0.386
  • 适合临床早期筛查,结果可解释性强

自动口吃检测系统在儿童语音中表现不佳,主要因发育中声音的高声学变异性,以及病理型口吃与正常发育性口吃之间的细微差别。我们提出Paediatric-HGNN,一种基于上下文感知整体-局部交互网络(CaPIN)的框架,专为儿童语音设计。不同于传统的1维信号建模,本方法构建异构图,捕捉词汇单元(词节点)与细粒度声学片段(帧节点)间的层级关系。在精心整理的儿科语料库UCLASS和FluencyBank上训练,该模型实现82.4%的加权准确率,典型口吃F1得分为0.386。通过建模词汇-声学层级交互,捕捉发育过程中的‘搜寻’行为特征,提供更鲁棒且可解释的早期临床干预工具。

原文摘要 · Abstract (English)

Automated stuttering detection (ASD) systems struggle with paediatric speech due to high acoustic variability in developing voices and the subtle distinction between pathological stuttering and typical developmental disfluencies. We introduce Paediatric-HGNN, a framework using a Context-aware Part-whole Interaction Network (CaPIN) tailored for paediatric data. Instead of conventional 1D signal modelling, our approach builds a heterogeneous graph capturing hierarchical relationships between lexical units (word nodes) and fine-grained acoustic segments (frame nodes). Trained on curated paediatric corpora (UCLASS and FluencyBank), Paediatric-HGNN achieves 82.4% weighted accuracy and a Typical Disfluency F1-score of 0.386. Modelling hierarchical lexical-acoustic interactions captures developmental "searching" behaviour, offering a more robust and interpretable tool for early clinical intervention.

语音识别儿童语音图神经网络口吃检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。