arXiv:2508.20805cs.CLcs.AI2025-08中稿 · APCIPA ASC 2025被引 2

用多模态数据对比机器学习与大模型检测抑郁效果

Exploring Machine Learning and Language Models for Multimodal Depression Detection

  • 融合音频、视频、文本特征,比较XGBoost、Transformer和大模型表现
  • 大模型在文本模态中识别抑郁信号更优,但跨模态融合仍有挑战
  • 适合关注心理健康智能诊断的研究者和临床应用开发者

本文针对首个多模态人格感知抑郁检测挑战赛,提出基于机器学习与深度学习的多模态抑郁检测方法。研究探索并对比了XGBoost、基于Transformer的架构以及大语言模型(LLMs)在音频、视频和文本特征上的表现。结果表明,各类模型在捕捉跨模态抑郁相关信号方面各有优势与局限,为心理健康预测中的有效多模态表征策略提供了重要参考。

原文摘要 · Abstract (English)

This paper presents our approach to the first Multimodal Personality-Aware Depression Detection Challenge, focusing on multimodal depression detection using machine learning and deep learning models. We explore and compare the performance of XGBoost, transformer-based architectures, and large language models (LLMs) on audio, video, and text features. Our results highlight the strengths and limitations of each type of model in capturing depression-related signals across modalities, offering insights into effective multimodal representation strategies for mental health prediction.

多模态抑郁症检测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。