arXiv:2409.08483cs.CLcs.AI2024-09被引 9

用文本摘要提升BERT模型对抑郁症状的识别准确率

A BERT-Based Summarization approach for depression detection

  • 用BERT模型结合文本摘要预处理,降低输入复杂度
  • 在DAIC-WOZ数据集上达0.67的测试F1分数,超越以往基准
  • 构建抑郁关键词词典,评估摘要质量与相关性

抑郁症是全球普遍的心理疾病,若未及时干预可能引发严重后果,尤其对反复发作者。早期干预可缓解症状,但实际应用面临挑战。利用机器学习和人工智能从多样化数据源自动检测抑郁指标具有前景,其中文本因能反映情绪、思维与感受而尤为关键。本研究基于虚拟访谈系统(如DAIC-WOZ数据集中的临床问卷),采用具备强大语义捕捉能力的BERT模型,将文本转化为数值表示,显著提升诊断精度。针对模型对长文本处理受限的问题,提出以文本摘要作为预处理手段,压缩输入长度与复杂度。在自研特征提取与分类框架中,测试集取得0.67的F1-score,超过所有先前基准;验证集达0.81,优于多数已有结果。同时构建抑郁词典,用于评估摘要质量与相关性,为后续研究提供重要工具。

原文摘要 · Abstract (English)

Depression is a globally prevalent mental disorder with potentially severe repercussions if not addressed, especially in individuals with recurrent episodes. Prior research has shown that early intervention has the potential to mitigate or alleviate symptoms of depression. However, implementing such interventions in a real-world setting may pose considerable challenges. A promising strategy involves leveraging machine learning and artificial intelligence to autonomously detect depression indicators from diverse data sources. One of the most widely available and informative data sources is text, which can reveal a person's mood, thoughts, and feelings. In this context, virtual agents programmed to conduct interviews using clinically validated questionnaires, such as those found in the DAIC-WOZ dataset, offer a robust means for depression detection through linguistic analysis. Utilizing BERT-based models, which are powerful and versatile yet use fewer resources than contemporary large language models, to convert text into numerical representations significantly enhances the precision of depression diagnosis. These models adeptly capture complex semantic and syntactic nuances, improving the detection accuracy of depressive symptoms. Given the inherent limitations of these models concerning text length, our study proposes text summarization as a preprocessing technique to diminish the length and intricacies of input texts. Implementing this method within our uniquely developed framework for feature extraction and classification yielded an F1-score of 0.67 on the test set surpassing all prior benchmarks and 0.81 on the validation set exceeding most previous results on the DAIC-WOZ dataset. Furthermore, we have devised a depression lexicon to assess summary quality and relevance. This lexicon constitutes a valuable asset for ongoing research in depression detection.

抑郁检测BERT文本摘要情感分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。