首个公开的德语阿尔茨海默病语音数据集,助力非侵入式早期诊断。
The PARLO Dementia Corpus: A German Multi-Center Resource for Alzheimer's Disease
- 跨九家诊所收集德语语音,含临床验证标签
- 包含8项标准化认知任务录音与转录,覆盖记忆与语言能力
- 适合研究神经退行性疾病语音分析、多模态模型的中文学者
阿尔茨海默病(AD)的早期可及性检测仍是重大挑战,当前诊断方法常依赖昂贵且侵入性的生物标志物。语音与语言分析作为无创、可扩展的潜在手段日益受到关注,但该领域研究受限于公开数据集匮乏,尤其缺乏非英语语种资源。本文介绍了帕尔洛痴呆语料库(PARLO Dementia Corpus, PDC),这是首个在德国九家学术记忆中心联合采集的、经临床验证的德语多中心数据集,涵盖伴有轻度认知障碍或轻中度痴呆的患者及认知健康对照者。语音通过八项标准化神经心理任务采集,包括命名、词汇流畅性、重复、图片描述、故事朗读与回忆等。数据集还包含人工校验的转录文本以及详细的受试者人口统计、临床和生物标志物信息。基准实验表明,自动语音识别、测试自动化评估及大语言模型分类均具备可行性,凸显回忆驱动型言语生产的诊断价值。PDC建立了首个公开可用的德语多模态与跨语言神经退行性疾病研究基准。
原文摘要 · Abstract (English)
Early and accessible detection of Alzheimer's disease (AD) remains a major challenge, as current diagnostic methods often rely on costly and invasive biomarkers. Speech and language analysis has emerged as a promising non-invasive and scalable approach to detecting cognitive impairment, but research in this area is hindered by the lack of publicly available datasets, especially for languages other than English. This paper introduces the PARLO Dementia Corpus (PDC), a new multi-center, clinically validated German resource for AD collected across nine academic memory clinics in Germany. The dataset comprises speech recordings from individuals with AD-related mild cognitive impairment and mild to moderate dementia, as well as cognitively healthy controls. Speech was elicited using a standardized test battery of eight neuropsychological tasks, including confrontation naming, verbal fluency, word repetition, picture description, story reading, and recall tasks. In addition to audio recordings, the dataset includes manually verified transcriptions and detailed demographic, clinical, and biomarker metadata. Baseline experiments on ASR benchmarking, automated test evaluation, and LLM-based classification illustrate the feasibility of automatic, speech-based cognitive assessment and highlight the diagnostic value of recall-driven speech production. The PDC thus establishes the first publicly available German benchmark for multi-modal and cross-lingual research on neurodegenerative diseases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。