TAMI通过时间对齐与缺失感知融合,提升老年轻度认知障碍者抑郁焦虑筛查精度。
TAMI: Temporally Aligned, Missingness-Aware, and Interpretable Multimodal Fusion for Mental Health Assessment in Older Adults with Mild Cognitive Impairment

- 按问答段落对齐多模态特征,动态编码模态缺失状态
- 抑郁与焦虑预测AUROC达0.68和0.69,时间对齐带来至少0.1性能提升
- 揭示开放性问题与眼动是抑郁关键信号,适合临床筛查流程设计
老年轻度认知障碍(MCI)患者常因医疗资源有限而未被及时诊断抑郁与焦虑。远程临床访谈的多模态分析虽具可扩展性,但现有方法存在三大局限:一是多模态特征因采样率不同导致时间错位,引发虚假跨模态关联;二是远程录音中模态缺失不均,但缺失值常被零填充,难以与真实低值区分;三是缺乏对模态、问题及访谈时刻的联合归因,限制细粒度临床解释。本文提出时间对齐、缺失感知、可解释的多模态融合框架TAMI。TAMI在49例老年MCI患者的访谈中,将语音、语言、面部与生理特征在问答段内对齐至统一时间轴,编码各模态随时间的缺失模式,并以问题上下文为条件进行融合。结果表明,抑郁与焦虑分类的受试者工作特征曲线下面积(AUROC)分别为0.68与0.69。细粒度时间对齐带来的性能增益最大(Δ≥0.1)。可解释性分析显示,抑郁依赖眼动与开放性问题,焦虑则依赖眼动与头部姿态,且归因在问题间分布均匀。仅使用开放性问题回答(5.1分钟)即可达到抑郁预测的AUROC 0.67,与完整访谈(19分钟)无显著差异(p>0.05)。研究支持以开放性问题为核心设计老年MCI抑郁筛查流程。
原文摘要 · Abstract (English)
Depression and anxiety in older adults with Mild Cognitive Impairment (MCI) are frequently underdiagnosed due to limited access to care. Multimodal analysis of remote clinical interviews is a scalable screening approach, but existing methods have three limitations. First, they do not correct temporal misalignment across multimodal features extracted at different resolutions, inducing spurious cross-modal associations. Second, remote recordings exhibit uneven modality dropout, but missing values are often zero-filled, making them indistinguishable from valid near-zero measurements. Finally, they do not jointly attribute predictions to modalities, questions, and interview moments, limiting fine-grained clinical interpretation. We propose a Temporally-Aligned, Missingness-Aware, Interpretable (TAMI) multimodal fusion framework. TAMI aligns speech, language, facial, and physiological features within question-answer segments on a shared timeline, encodes modality-level missingness over time, and conditions fusion on question context. In interviews with 49 older adults with MCI, TAMI achieved area under the receiver operating characteristic curve (AUROC) scores of 0.68 (depression) and 0.69 (anxiety). Fine-grained temporal alignment of multimodal features produced the largest performance gain ($Δ{\geq}0.1$). Multi-level interpretability analysis revealed that depression classification relied on eyegaze and open-ended questions, while anxiety classification depended on eyegaze and head pose, with attribution uniformly distributed across questions. Using only responses to the open-ended questions (5.1min), the depression model achieved an AUROC score of 0.67, which was not significantly different from using the full interview (19min) ($p>0.05$). Our findings support designing interview protocols centered on open-ended questions for depression screening in older adults with MCI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。