用对比学习提升多语言多图描述的轻度认知障碍检测效果
Unveil Multi-Picture Descriptions for Multilingual Mild Cognitive Impairment Detection via Contrastive Learning
- 通过对比学习增强文本和图像的表征能力
- 多模态融合使轻度认知障碍检测的召回率提升至75.2%
- 适合关注跨语言医疗诊断与多模态模型的研究者
从图片描述中检测轻度认知障碍具有重要意义但挑战重重,尤其在多语言和多图场景下。以往研究主要聚焦于英语母语者对单张图片(如‘饼干盗窃图’)的描述。TAUKDIAL-2024挑战赛拓展了这一范围,引入多语言描述者与多张图片,带来新的内容依赖性分析难题。为此,我们提出一个三组件框架:(1) 通过有监督对比学习增强判别性表征;(2) 引入图像模态,而非仅依赖语音与文本;(3) 采用专家乘积(PoE)策略缓解虚假相关与过拟合问题。相较于仅使用文本的基线模型,本框架在未加权平均召回率(UAR)上提升7.1%(从68.1%到75.2%),F1分数提升2.9%(从80.6%到83.5%)。值得注意的是,对比学习对文本模态的增益高于语音模态。结果表明该框架在多语言、多图场景下的有效性。
原文摘要 · Abstract (English)
Detecting Mild Cognitive Impairment from picture descriptions is critical yet challenging, especially in multilingual and multiple picture settings. Prior work has primarily focused on English speakers describing a single picture (e.g., the 'Cookie Theft'). The TAUKDIAL-2024 challenge expands this scope by introducing multilingual speakers and multiple pictures, which presents new challenges in analyzing picture-dependent content. To address these challenges, we propose a framework with three components: (1) enhancing discriminative representation learning via supervised contrastive learning, (2) involving image modality rather than relying solely on speech and text modalities, and (3) applying a Product of Experts (PoE) strategy to mitigate spurious correlations and overfitting. Our framework improves MCI detection performance, achieving a +7.1% increase in Unweighted Average Recall (UAR) (from 68.1% to 75.2%) and a +2.9% increase in F1 score (from 80.6% to 83.5%) compared to the text unimodal baseline. Notably, the contrastive learning component yields greater gains for the text modality compared to speech. These results highlight our framework's effectiveness in multilingual and multi-picture MCI detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。