用Mamba动态融合多模态数据,提升抑郁症检测准确率
CAF-Mamba: Mamba-Based Cross-Modal Adaptive Attention Fusion for Multimodal Depression Detection
- 基于Mamba构建跨模态自适应注意力融合机制
- 在LMVD和D-Vlog数据集上超越现有方法
- 适合做多模态情感分析与心理健康检测的研究者
抑郁症是一种普遍的心理健康障碍,严重损害日常生活功能和生活质量。尽管近期深度学习方法在抑郁症检测中展现出潜力,但多数方法仅使用有限特征类型,忽略显式跨模态交互,且采用简单的拼接或静态加权进行融合。为克服这些局限,我们提出CAF-Mamba,一种基于Mamba的跨模态自适应注意力融合框架。CAF-Mamba不仅能显式和隐式地捕捉跨模态交互,还可通过模态级注意力机制动态调整各模态贡献,实现更有效的多模态融合。在两个真实场景基准数据集LMVD和D-Vlog上的实验表明,CAF-Mamba持续优于现有方法,达到最先进性能。代码已开源:https://github.com/zbw-zhou/CAF-Mamba。
原文摘要 · Abstract (English)
Depression is a prevalent mental health disorder that severely impairs daily functioning and quality of life. While recent deep learning approaches for depression detection have shown promise, most rely on limited feature types, overlook explicit cross-modal interactions, and employ simple concatenation or static weighting for fusion. To overcome these limitations, we propose CAF-Mamba, a novel Mamba-based cross-modal adaptive attention fusion framework. CAF-Mamba not only captures cross-modal interactions explicitly and implicitly, but also dynamically adjusts modality contributions through a modality-wise attention mechanism, enabling more effective multimodal fusion. Experiments on two in-the-wild benchmark datasets, LMVD and D-Vlog, demonstrate that CAF-Mamba consistently outperforms existing methods and achieves state-of-the-art performance. Our code is available at https://github.com/zbw-zhou/CAF-Mamba.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。