arXiv:2511.01892cs.LGcs.CL2025-11中稿 · IEEE EMBC 2025被引 2

用检索增强生成提升抑郁检测的多模态表现

Retrieval-Augmented Multimodal Depression Detection

  • 从情感数据集检索相关语义内容,生成情绪提示作为辅助模态
  • 在AVEC 2019上达成CCC 0.593、MAE 3.95的领先效果
  • 适合关注情绪理解与可解释性的多模态抑郁研究者

多模态深度学习通过融合文本、音频和视频信号,在抑郁检测中展现出潜力。近期工作利用情感分析增强情绪理解,但存在计算成本高、领域不匹配和知识静态等问题。为此,我们提出一种新型检索增强生成(RAG)框架:给定抑郁相关文本,从情感数据集中检索语义相关的感情内容,并使用大语言模型(LLM)生成情绪提示作为辅助模态。该提示丰富了情绪表征并提升可解释性。在AVEC 2019数据集上的实验表明,该方法达到当前最优性能,一致性相关系数(CCC)为0.593,平均绝对误差(MAE)为3.95,优于以往的迁移学习与多任务学习基线。

原文摘要 · Abstract (English)

Multimodal deep learning has shown promise in depression detection by integrating text, audio, and video signals. Recent work leverages sentiment analysis to enhance emotional understanding, yet suffers from high computational cost, domain mismatch, and static knowledge limitations. To address these issues, we propose a novel Retrieval-Augmented Generation (RAG) framework. Given a depression-related text, our method retrieves semantically relevant emotional content from a sentiment dataset and uses a Large Language Model (LLM) to generate an Emotion Prompt as an auxiliary modality. This prompt enriches emotional representation and improves interpretability. Experiments on the AVEC 2019 dataset show our approach achieves state-of-the-art performance with CCC of 0.593 and MAE of 3.95, surpassing previous transfer learning and multi-task learning baselines.

抑郁检测多模态RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。