通过解耦与自适应融合,提升抑郁诊断的多模态分析准确性
IDRL: An Individual-Aware Multimodal Depression-Related Representation Learning Framework for Depression Diagnosis
- 分离模态共性、特性与无关信息,增强跨模态对齐
- 动态调整个体差异特征权重,实现个性化融合
- 适合临床辅助诊断与个性化心理健康评估
抑郁症是一种严重的精神障碍,可靠识别对早期干预和治疗至关重要。多模态抑郁检测通过联合建模多种模态的互补信息以提升诊断性能。然而,现有方法存在两大问题:1)模态间不一致及非抑郁相关干扰,即抑郁线索在不同模态中可能冲突,大量无关内容掩盖关键抑郁信号;2)个体抑郁表现差异大,导致各模态与特征重要性不一,影响可靠融合。为此,本文提出个体感知的多模态抑郁相关表征学习框架(IDRL)。IDRL 1)将多模态表征解耦为模态共有的抑郁空间、模态特异的抑郁空间和非抑郁相关空间,增强模态对齐并抑制无关信息;2)引入个体感知的模态融合模块(IAF),根据特征预测价值动态调整解耦后抑郁相关特征的权重,实现针对不同个体的自适应跨模态融合。大量实验表明,IDRL在多模态抑郁检测任务中取得了更优且鲁棒的性能。
原文摘要 · Abstract (English)
Depression is a severe mental disorder, and reliable identification plays a critical role in early intervention and treatment. Multimodal depression detection aims to improve diagnostic performance by jointly modeling complementary information from multiple modalities. Recently, numerous multimodal learning approaches have been proposed for depression analysis; however, these methods suffer from the following limitations: 1) inter-modal inconsistency and depression-unrelated interference, where depression-related cues may conflict across modalities while substantial irrelevant content obscures critical depressive signals, and 2) diverse individual depressive presentations, leading to individual differences in modality and cue importance that hinder reliable fusion. To address these issues, we propose Individual-aware Multimodal Depression-related Representation Learning Framework (IDRL) for robust depression diagnosis. Specifically, IDRL 1) disentangles multimodal representations into a modality-common depression space, a modality-specific depression space, and a depression-unrelated space to enhance modality alignment while suppressing irrelevant information, and 2) introduces an individual-aware modality-fusion module (IAF) that dynamically adjusts the weights of disentangled depression-related features based on their predictive significance, thereby achieving adaptive cross-modal fusion for different individuals. Extensive experiments demonstrate that IDRL achieves superior and robust performance for multimodal depression detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。