arXiv:2512.06447cs.CV2025-12被引 9

提出统一框架,让抑郁识别模型在数据缺失时仍稳定有效。

Towards Stable Cross-Domain Depression Recognition under Missing Modalities

  • 用掩码和提示词统一处理多源异构数据输入
  • 自适应融合音频视频特征,支持部分数据缺失场景
  • 在5个真实数据集上表现优于现有模型,适合实际应用

抑郁症带来严重公共健康风险,亟需及时、可扩展的筛查手段。多模态自动抑郁检测(ADD)具前景,但现有基于音频和视频的方法缺乏统一通用框架,对缺失模态稳定性差,而真实数据中缺失现象普遍。本文提出基于多模态大语言模型的稳定跨域抑郁识别统一框架(SCD-MLLM)。该框架支持异构数据整合与处理,并在模态不全时保持稳定。核心包含:(i) 多源数据输入适配器(MDIA),通过掩码机制和任务特定提示将不同来源的抑郁相关输入转化为统一标记序列,解决数据不一致问题;(ii) 模态感知自适应融合模块(MAFM),通过共享投影机制自适应融合音视频特征,提升缺失模态下的鲁棒性。在五个公开异构抑郁数据集(CMDC、AVEC2014、DAIC-WOZ、DVlog、EATD)的多数据集联合训练设置下进行实验,无论完整还是部分模态条件下,SCD-MLLM均超越当前最优模型及主流商用大模型(Gemini、GPT),展现出更强跨域泛化能力、更优多模态线索捕捉能力以及在真实应用中对缺失模态的强稳定性。

原文摘要 · Abstract (English)

Depression poses serious public health risks, including suicide, underscoring the urgency of timely and scalable screening. Multimodal automatic depression detection (ADD) offers a promising solution; however, widely studied audio- and video-based ADD methods lack a unified, generalizable framework for diverse depression recognition scenarios and show limited stability to missing modalities, which are common in real-world data. In this work, we propose a unified framework for Stable Cross-Domain Depression Recognition based on Multimodal Large Language Model (SCD-MLLM). The framework supports the integration and processing of heterogeneous depression-related data collected from varied sources while maintaining stability in the presence of incomplete modality inputs. Specifically, SCD-MLLM introduces two key components: (i) Multi-Source Data Input Adapter (MDIA), which employs masking mechanism and task-specific prompts to transform heterogeneous depression-related inputs into uniform token sequences, addressing inconsistency across diverse data sources; (ii) Modality-Aware Adaptive Fusion Module (MAFM), which adaptively integrates audio and visual features via a shared projection mechanism, enhancing resilience under missing modality conditions. e conduct comprehensive experiments under multi-dataset joint training settings on five publicly available and heterogeneous depression datasets from diverse scenarios: CMDC, AVEC2014, DAIC-WOZ, DVlog, and EATD. Across both complete and partial modality settings, SCD-MLLM outperforms state-of-the-art (SOTA) models as well as leading commercial LLMs (Gemini and GPT), demonstrating superior cross-domain generalization, enhanced ability to capture multimodal cues of depression, and strong stability to missing modality cases in real-world applications.

抑郁识别多模态大模型缺失数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。