融合眼动与面部特征,提升阿尔茨海默病早期诊断准确率
Cross-Enhanced Multimodal Fusion of Eye-Tracking and Facial Features for Alzheimer's Disease Diagnosis
- 设计交叉增强注意力模块,动态捕捉眼动与面部信息的相互作用
- 在25例患者与25例健康人数据上达到95.11%分类准确率
- 适合关注多模态融合与神经退行性疾病智能诊断的研究者
阿尔茨海默病(AD)的精准诊断对及时干预和延缓疾病进展至关重要。多模态诊断方法通过整合行为与感知领域的互补信息展现出巨大潜力。眼动追踪和面部特征是反映注意力分布与神经认知状态的重要指标。然而,极少研究探索二者联合用于辅助诊断。本文提出一种多模态交叉增强融合框架,协同利用眼动与面部特征进行AD检测。框架包含两个关键模块:(a) 交叉增强融合注意力模块(CEFAM),通过交叉注意力与全局增强建模跨模态交互;(b) 方向感知卷积模块(DACM),利用水平-垂直感受野捕获精细方向性面部特征。两者共同实现自适应、判别性强的多模态表征学习。为支持本研究,构建了一个同步多模态数据集,包含25名AD患者与25名健康对照(HC),在视觉记忆搜索范式中记录对齐的面部视频与眼动序列,提供生态效度高的评估资源。大量实验表明,该框架优于传统后期融合与特征拼接方法,在区分AD与HC任务中达到95.11%的分类准确率,凸显其在显式建模跨模态依赖关系与模态特异性贡献方面的优越鲁棒性与诊断性能。
原文摘要 · Abstract (English)
Accurate diagnosis of Alzheimer's disease (AD) is essential for enabling timely intervention and slowing disease progression. Multimodal diagnostic approaches offer considerable promise by integrating complementary information across behavioral and perceptual domains. Eye-tracking and facial features, in particular, are important indicators of cognitive function, reflecting attentional distribution and neurocognitive state. However, few studies have explored their joint integration for auxiliary AD diagnosis. In this study, we propose a multimodal cross-enhanced fusion framework that synergistically leverages eye-tracking and facial features for AD detection. The framework incorporates two key modules: (a) a Cross-Enhanced Fusion Attention Module (CEFAM), which models inter-modal interactions through cross-attention and global enhancement, and (b) a Direction-Aware Convolution Module (DACM), which captures fine-grained directional facial features via horizontal-vertical receptive fields. Together, these modules enable adaptive and discriminative multimodal representation learning. To support this work, we constructed a synchronized multimodal dataset, including 25 patients with AD and 25 healthy controls (HC), by recording aligned facial video and eye-tracking sequences during a visual memory-search paradigm, providing an ecologically valid resource for evaluating integration strategies. Extensive experiments on this dataset demonstrate that our framework outperforms traditional late fusion and feature concatenation methods, achieving a classification accuracy of 95.11% in distinguishing AD from HC, highlighting superior robustness and diagnostic performance by explicitly modeling inter-modal dependencies and modality-specific contributions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。