区分说话中的正常与异常不流畅,提升语音助手对口吃者识别准确率
Typical vs. Atypical Disfluency Classification: Introducing the IIITH-TISA Corpus and Temporal Context-Based Feature Representations
- 用时序特征捕捉语音中局部与全局时间上下文
- 在自建语料上实现85.01%平均F1分数,优于传统方法
- 适合语音助手机器人优化与儿童口吃早期筛查研究
自发性口语中的言语不流畅可分为正常与异常两类。正常不流畅如停顿、重复是日常交流的自然现象,而异常不流畅则可能提示口吃等病理状况。准确区分两者对提升语音助手对口吃人群(PWS)的服务体验至关重要,可避免因误判语句结束导致的提前中断。同时,该技术有助于儿童口吃早期发现,防止误诊为语言发展性不流畅。本研究提出首个印度英语口吃语料库IIITH-TISA,并扩展了IIITH-IED数据集以标注正常不流畅。我们采用感知增强的零时窗倒谱系数(PE-ZTWCC)与移位差分倒谱(SDC)作为浅层时延神经网络(TDNN)的输入特征,有效捕捉局部与广泛的时间上下文信息。实验表明,该方法在不流畅分类任务上达到85.01%的平均F1得分,显著优于传统特征。
原文摘要 · Abstract (English)
Speech disfluencies in spontaneous communication can be categorized as either typical or atypical. Typical disfluencies, such as hesitations and repetitions, are natural occurrences in everyday speech, while atypical disfluencies are indicative of pathological disorders like stuttering. Distinguishing between these categories is crucial for improving voice assistants (VAs) for Persons Who Stutter (PWS), who often face premature cutoffs due to misidentification of speech termination. Accurate classification also aids in detecting stuttering early in children, preventing misdiagnosis as language development disfluency. This research introduces the IIITH-TISA dataset, the first Indian English stammer corpus, capturing atypical disfluencies. Additionally, we extend the IIITH-IED dataset with detailed annotations for typical disfluencies. We propose Perceptually Enhanced Zero-Time Windowed Cepstral Coefficients (PE-ZTWCC) combined with Shifted Delta Cepstra (SDC) as input features to a shallow Time Delay Neural Network (TDNN) classifier, capturing both local and wider temporal contexts. Our method achieves an average F1 score of 85.01% for disfluency classification, outperforming traditional features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。