arXiv:2502.19208cs.CL2025-02被引 7

构建首个多语言对话数据集,助力阿尔茨海默病早期检测

MultiConAD: A Unified Multilingual Conversational Dataset for Early Alzheimer's Detection

  • 整合16个公开数据集,覆盖英、西、中、希腊语的语音与文本对话
  • 实现从正常到轻度认知障碍的细粒度分类,提升早期干预可能
  • 发现多语言训练对部分语言有效,为跨语言模型设计提供依据

痴呆症是一种进行性认知综合征,阿尔茨海默病(AD)是其主要病因。基于对话的AD检测提供了低成本替代方案,因语言功能障碍是AD的早期生物标志物。然而,以往研究多将AD检测视为二分类问题,难以识别轻度认知障碍(MCI)这一关键早期阶段;且研究主要依赖英文单语数据集,限制了跨语言泛化能力。为此,本文有三项贡献:第一,首次构建一个统一的多语言对话数据集,整合16个公开的痴呆相关对话数据集,涵盖英语、西班牙语、中文和希腊语,包含多种认知评估任务的音频与文本数据;第二,采用更细粒度的分类任务(包括正常、MCI等),评估稀疏与密集文本表示下的各类分类器性能;第三,在单语和多语言设置下开展实验,发现部分语言在多语言训练中表现更好,而其他语言则独立训练更优。该研究揭示了多语言AD检测的挑战,推动未来针对语言特异性方法及提升模型泛化与鲁棒性的研究。

原文摘要 · Abstract (English)

Dementia is a progressive cognitive syndrome with Alzheimer's disease (AD) as the leading cause. Conversation-based AD detection offers a cost-effective alternative to clinical methods, as language dysfunction is an early biomarker of AD. However, most prior research has framed AD detection as a binary classification problem, limiting the ability to identify Mild Cognitive Impairment (MCI)-a crucial stage for early intervention. Also, studies primarily rely on single-language datasets, mainly in English, restricting cross-language generalizability. To address this gap, we make three key contributions. First, we introduce a novel, multilingual dataset for AD detection by unifying 16 publicly available dementia-related conversational datasets. This corpus spans English, Spanish, Chinese, and Greek and incorporates both audio and text data derived from a variety of cognitive assessment tasks. Second, we perform finer-grained classification, including MCI, and evaluate various classifiers using sparse and dense text representations. Third, we conduct experiments in monolingual and multilingual settings, finding that some languages benefit from multilingual training while others perform better independently. This study highlights the challenges in multilingual AD detection and enables future research on both language-specific approaches and techniques aimed at improving model generalization and robustness.

阿尔茨海默病多语言对话分析早期检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。