arXiv:2503.01226q-bio.NCcs.LG2025-03被引 1

用文本和音频的上下文嵌入,提升阿尔茨海默病早期检测准确率。

Dementia Insights: A Context-Based MultiModal Approach

  • 融合GPT文本与CLAP音频的上下文特征进行多模态分析
  • F1得分达83.33%,优于现有最先进模型
  • 无需专家标注,直接使用原始文本效果更优

阿尔茨海默病是一种进行性神经退行性疾病,影响记忆、推理和日常功能,给个人和医疗系统带来挑战。早期检测对及时干预、延缓疾病进展至关重要。大型预训练模型(如GPT、BERT、CLAP)在识别认知障碍方面展现出潜力。但现有研究多依赖专家标注数据集且采用单模态方法,限制了鲁棒性和可扩展性。本文提出一种基于上下文的多模态方法,分别利用各模态表现最优的LPM提取文本与音频特征,并引入上下文嵌入提升检测性能。此外,探索了上下文学习(ICL)作为补充技术。结果表明,结合GPT文本嵌入与CLAP音频特征的融合模型,实现83.33%的F1得分,超越当前最优模型。更重要的是,原始文本数据的表现优于专家标注数据集,证明LPM能无需大量人工标注即可提取有意义的语言与声学模式。该研究为低成本、非侵入式诊断工具提供了可行路径,推动个性化认知健康监测发展。

原文摘要 · Abstract (English)

Dementia, a progressive neurodegenerative disorder, affects memory, reasoning, and daily functioning, creating challenges for individuals and healthcare systems. Early detection is crucial for timely interventions that may slow disease progression. Large pre-trained models (LPMs) for text and audio, such as Generative Pre-trained Transformer (GPT), Bidirectional Encoder Representations from Transformers (BERT), and Contrastive Language-Audio Pretraining (CLAP), have shown promise in identifying cognitive impairments. However, existing studies generally rely heavily on expert-annotated datasets and unimodal approaches, limiting robustness and scalability. This study proposes a context-based multimodal method, integrating both text and audio data using the best-performing LPMs in each modality. By incorporating contextual embeddings, our method improves dementia detection performance. Additionally, motivated by the effectiveness of contextual embeddings, we further experimented with a context-based In-Context Learning (ICL) as a complementary technique. Results show that GPT-based embeddings, particularly when fused with CLAP audio features, achieve an F1-score of $83.33\%$, surpassing state-of-the-art dementia detection models. Furthermore, raw text data outperforms expert-annotated datasets, demonstrating that LPMs can extract meaningful linguistic and acoustic patterns without extensive manual labeling. These findings highlight the potential for scalable, non-invasive diagnostic tools that reduce reliance on costly annotations while maintaining high accuracy. By integrating multimodal learning with contextual embeddings, this work lays the foundation for future advancements in personalized dementia detection and cognitive health research.

阿尔茨海默病多模态学习上下文嵌入零样本检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。