比较语音输入方式对帕金森病零样本检测的影响
Zero-Shot Parkinson's Disease Detection from Speech: Comparing Large Audio and Language Models

- 用手工特征与原始波形两种输入,让大模型判断帕金森病
- 低资源语言中手工特征更稳定,原始音频在特定数据集有提升
- 适合医疗AI、语音诊断研究者参考
大型音频和语言模型最近在多个领域展现出零样本推理能力。然而,语音输入形式(手工提取的声学特征或原始音频波形)如何影响帕金森病(PD)检测性能,尤其是在不同语言下的表现,仍不明确。本研究系统比较了两种零样本PD检测输入模态:(i) 由通用大语言模型分析的语音记录中提取的手工声学特征;(ii) 由音频模型直接处理的原始波形。在四种语言的帕金森病语音数据集上实验发现,性能随输入模态、语音任务和语言而异。在低资源语言(如孟加拉语)中,手工特征表现更稳定;而原始音频输入在部分数据集上带来显著增益。结果凸显了输入模态对零样本语音帕金森病检测的关键影响。
原文摘要 · Abstract (English)
Large audio and language models have recently demonstrated zero-shot reasoning capabilities across various domains. However, it remains unclear how the form of audio input, whether handcrafted acoustic features extracted from speech or the raw audio waveform itself, affects performance for Parkinson's disease (PD) detection across different languages. In this study, we systematically compare two input modalities for zero-shot PD detection: (i) handcrafted acoustic features extracted from speech recordings analyzed by a general-purpose LLM, and (ii) direct waveform input analyzed by audio-capable models. Experiments on PD speech datasets in four languages show that performance varies across input modalities, speech tasks, and languages. Handcrafted acoustic features provide more stable performance in a low-resource language (e.g., Bengali), whereas audio input yields dataset-dependent gains. These findings highlight the impact of input modality on zero-shot PD detection from speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。