用语音表征模型提升儿童语言障碍分类,效果优于大模型。
Multimodal LLMs are not all you need for Pediatric Speech Language Pathology

- 分层分类:从二分类到类型与症状逐级判断。
- 语音表征模型在所有任务上显著超越大模型,误差率降低37%。
- 专为儿科语音病理设计数据增强,缓解偏差问题,适合临床研究者使用。
言语发音障碍(SSD)影响约5%的儿童,但言语治疗师严重缺编且工作负荷过重。我们基于细粒度多任务SLPHelmUltraSuitePlus基准,测试了一种分层分类方法,依次进行二分类、类型与症状分类。通过微调语音表征模型(SRM)并采用针对性数据增强,缓解了先前研究发现的偏差问题,在该基准所有临床任务中均取得更好表现。同时,我们的数据增强方法也提升了自动语音识别(ASR)性能。结果表明,SRM在所有评估任务中均大幅优于基于大语言模型(LLM)的现有方法。我们已公开模型与代码,以推动后续研究。
原文摘要 · Abstract (English)
Speech Sound Disorders (SSD) affect roughly five percent of children, yet speech-language pathologists face severe staffing shortages and unmanageable caseloads. We test a hierarchical approach to SSD classification on the granular multi-task SLPHelmUltraSuitePlus benchmark. We propose a cascading approach from binary classification to type, and symptom classification. By fine-tuning Speech Representation Models (SRM), and using targeted data augmentation we mitigate biases found by previous works, and improve upon all clinical tasks in the benchmark. We also treat Automatic Speech Recognition (ASR) with our data augmentation approach. Our results demonstrate that SRM consistently outperform the LLM-based state-of-the-art across all evaluated tasks by a large margin. We publish our models and code to foster future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。