arXiv:2510.14922cs.AIcs.CL2025-10

融合脑电、语音与文本三模态数据,提升抑郁症自动检测效果

TRI-DEP: A Trimodal Comparative Study for Depression Detection Using Speech, Text, and EEG

  • 对比手写特征与预训练嵌入,评估多模态编码器与融合策略
  • 三模态联合模型在独立受试者测试中达最新性能纪录
  • 为抑郁症多模态分析提供可复现的基准方法,适合临床辅助研究

抑郁症是一种普遍存在的心理健康障碍,其自动化检测仍具挑战性。以往研究多采用单模态或双模态方法,虽显示一定潜力,但存在范围局限、特征对比不系统、评估协议不一致等问题。本文系统探索了脑电(EEG)、语音与文本三模态在特征表示和建模策略上的表现,对比手工特征与预训练嵌入,评估不同神经编码器效果,比较单、双、三模态配置及融合方式,并重点分析脑电的作用。所有实验均采用一致的受试者独立划分,确保结果稳健可复现。结果表明:(i) EEG、语音与文本三者结合显著提升检测性能;(ii) 预训练嵌入优于手工特征;(iii) 经精心设计的三模态模型达到当前最优水平。本工作为后续多模态抑郁症检测研究奠定基础。

原文摘要 · Abstract (English)

Depression is a widespread mental health disorder, yet its automatic detection remains challenging. Prior work has explored unimodal and multimodal approaches, with multimodal systems showing promise by leveraging complementary signals. However, existing studies are limited in scope, lack systematic comparisons of features, and suffer from inconsistent evaluation protocols. We address these gaps by systematically exploring feature representations and modelling strategies across EEG, together with speech and text. We evaluate handcrafted features versus pre-trained embeddings, assess the effectiveness of different neural encoders, compare unimodal, bimodal, and trimodal configurations, and analyse fusion strategies with attention to the role of EEG. Consistent subject-independent splits are applied to ensure robust, reproducible benchmarking. Our results show that (i) the combination of EEG, speech and text modalities enhances multimodal detection, (ii) pretrained embeddings outperform handcrafted features, and (iii) carefully designed trimodal models achieve state-of-the-art performance. Our work lays the groundwork for future research in multimodal depression detection.

抑郁症检测多模态学习脑电分析语音识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。