机器比医生更准识别帕金森语音,尤其对年轻、轻症和女性患者。
Comparison of Speech Tasks in Human Expert and Machine Detection of Parkinson's Disease
- 用Whisper模型分析五类语音任务,仅靠音频判断帕金森病。
- 在年轻、轻症及女性患者中,模型准确率超越人类专家。
- 适合临床辅助诊断与语音健康监测场景使用。
帕金森病患者的语音包含疾病存在与进展的重要线索。本文研究人类专家基于五类语音任务(发声、句子复述、朗读、回忆、图片描述)判断疾病存在性的依据。通过听觉测试评估临床医生仅凭音频识别帕金森病的准确性,并利用Whisper模型进行机器学习检测实验。结果显示,在仅提供音频的情况下,Whisper在多数任务中的表现达到或优于人类专家,尤其在年轻患者、轻症患者和女性患者等挑战性子群体中表现突出。模型对复杂情况下声学特征的识别能力,可弥补人类专家在多模态与情境理解方面的优势。
原文摘要 · Abstract (English)
The speech of people with Parkinson's Disease (PD) has been shown to hold important clues about the presence and progression of the disease. We investigate the factors based on which humans experts make judgments of the presence of disease in speech samples over five different speech tasks: phonations, sentence repetition, reading, recall, and picture description. We make comparisons by conducting listening tests to determine clinicians accuracy at recognizing signs of PD from audio alone, and we conduct experiments with a machine learning system for detection based on Whisper. Across tasks, Whisper performs on par or better than human experts when only audio is available, especially on challenging but important subgroups of the data: younger patients, mild cases, and female patients. Whisper's ability to recognize acoustic cues in difficult cases complements the multimodal and contextual strengths of human experts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。