语言特征对帕金森病语音早期检测至关重要,文本模型表现不逊于声学模型。
Does Language Matter for Early Detection of Parkinson's Disease from Speech?
- 用预训练模型对比不同数据类型,验证语言信息价值
- 纯文本模型性能媲美声学特征模型,多语言Whisper更优
- 音频集预训练提升持续发音任务效果,但对自发语音无效
利用语音样本作为帕金森病(PD)的生物标志物具有广阔前景,但数据采集与分析方法在文献中存在较大争议。早期研究多采用持续元音发声(SVP)任务,而近年也探索了更具认知负荷的任务录音。为评估语言在PD检测中的作用,我们测试了不同数据类型和预训练目标的预训练模型,发现:(1)仅使用文本的模型性能与仅使用声学特征的模型相当;(2)多语言Whisper模型优于自监督模型,而单语言Whisper表现更差;(3)AudioSet预训练能提升SVP任务表现,但对自发语音无益。这些结果共同凸显语言在帕金森病早期检测中的关键作用。
原文摘要 · Abstract (English)
Using speech samples as a biomarker is a promising avenue for detecting and monitoring the progression of Parkinson's disease (PD), but there is considerable disagreement in the literature about how best to collect and analyze such data. Early research in detecting PD from speech used a sustained vowel phonation (SVP) task, while some recent research has explored recordings of more cognitively demanding tasks. To assess the role of language in PD detection, we tested pretrained models with varying data types and pretraining objectives and found that (1) text-only models match the performance of vocal-feature models, (2) multilingual Whisper outperforms self-supervised models whereas monolingual Whisper does worse, and (3) AudioSet pretraining improves performance on SVP but not spontaneous speech. These findings together highlight the critical role of language for the early detection of Parkinson's disease.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。