用多模态数据对比机器学习与大模型检测抑郁效果
Exploring Machine Learning and Language Models for Multimodal Depression Detection
- 融合音频、视频、文本特征,比较XGBoost、Transformer和大模型表现
- 大模型在文本模态中识别抑郁信号更优,但跨模态融合仍有挑战
- 适合关注心理健康智能诊断的研究者和临床应用开发者
本文针对首个多模态人格感知抑郁检测挑战赛,提出基于机器学习与深度学习的多模态抑郁检测方法。研究探索并对比了XGBoost、基于Transformer的架构以及大语言模型(LLMs)在音频、视频和文本特征上的表现。结果表明,各类模型在捕捉跨模态抑郁相关信号方面各有优势与局限,为心理健康预测中的有效多模态表征策略提供了重要参考。
原文摘要 · Abstract (English)
This paper presents our approach to the first Multimodal Personality-Aware Depression Detection Challenge, focusing on multimodal depression detection using machine learning and deep learning models. We explore and compare the performance of XGBoost, transformer-based architectures, and large language models (LLMs) on audio, video, and text features. Our results highlight the strengths and limitations of each type of model in capturing depression-related signals across modalities, offering insights into effective multimodal representation strategies for mental health prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。