arXiv:2605.14888cs.SDcs.LG2026-05被引 1

构建大规模语音数据集,助力早期认知衰退自动检测

PROCESS-2: A Benchmark Speech Corpus for Early Cognitive Impairment Detection

论文配图:PROCESS-2: A Benchmark Speech Corpus for Early Cognitive Impairment Detection
图 1 · 摘自论文原文
  • 基于真实场景采集200名健康者、150名轻度痴呆、50名痴呆患者的语音数据
  • 包含约21小时语音,支持自发与任务导向性语言分析,分好训练测试集
  • 数据经临床验证,可复现基准模型性能,适合医疗AI研究者使用

基于语音的分析为检测认知衰退提供了一种可扩展且无创的方法,但受限于在真实条件下收集的经临床验证数据集稀缺。我们推出PROCESS-2,一个大规模语音数据集,旨在支持从自发和任务导向语音中自动评估认知障碍的研究。该数据集包含200名健康对照者、150名轻度认知障碍患者和50名痴呆诊断者,通过CognoMemory数字评估平台采集。每位参与者完成一次评估会话,包括图片描述和言语流畅性任务,附有手动校验的转录文本和个体层面元数据。PROCESS-2包含约21小时语音音频,并设有预定义的训练/测试划分。全面的技术验证评估了人口统计平衡性、临床一致性、录音稳定性、嵌入空间结构及可复现的基线建模表现,展示了具有临床意义的组间分离性,并在不同建模方法下保持稳定性能,同时保留真实对话中的变异性。PROCESS-2通过Hugging Face以受控访问方式发布,确保负责任复用并保护参与者隐私,为语音驱动的认知评估研究提供可复现的基准资源。

原文摘要 · Abstract (English)

Speech-based analysis offers a scalable and non-invasive approach for detecting cognitive decline, yet progress has been constrained by the limited availability of clinically validated datasets collected under realistic conditions. We introduce PROCESS-2, a large-scale speech dataset designed to support research on automatic assessment of cognitive impairment from spontaneous and task-oriented speech. The dataset comprises recordings from 200 healthy controls, 150 mild cognitive impairment, and 50 dementia diagnoses collected using the CognoMemory digital assessment platform. Each participant completed a single assessment session, including picture description and verbal fluency tasks, accompanied by manually verified transcripts and participant-level metadata. PROCESS-2 contains approximately 21 hours of speech audio with predefined train/test partitions. Comprehensive technical validation evaluated demographic balance, clinical consistency, recording stability, embedding-space structure, and reproducible baseline modelling performance, demonstrating clinically meaningful group separation and stable performance across modelling approaches while preserving real-world conversational variability. PROCESS-2 is released under controlled access via Hugging Face to enable responsible reuse while protecting participant privacy, providing a reproducible benchmark resource for speech-based cognitive assessment research.

语音分析认知障碍数据集医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。