arXiv:2412.01353cs.HCcs.AI2024-12中稿 · IEEE International…被引 2

用小模型分析社交媒体,高效识别自杀风险

Su-RoBERTa: A Semi-supervised Approach to Predicting Suicide Risk through Social Media using Base Language Models

  • 基于鲁棒的基线模型,结合有标签与无标签数据微调
  • 在Reddit数据上达69.84%加权F1分数,优于大模型
  • 适合资源有限但需心理风险筛查的场景

近年来,越来越多用户在社交媒体上表达心理状态。利用这些数据,可构建AI系统以评估个体心理健康状况,如自杀风险。本文基于Reddit数据,采用基础语言模型进行自杀风险预测研究。我们证明,参数量小于500M的小型语言模型同样有效,优于参数量大于500M的大模型。提出Su-RoBERTa模型,在自杀风险预测任务上对RoBERTa进行微调,融合标注与未标注的Reddit数据,并通过GPT-2进行数据增强以缓解类别不平衡问题。最终评估中,该模型获得69.84%的加权F1分数。本研究展示了基础语言模型在心理健康风险因素分析中的有效性,具备高效的计算流程。

原文摘要 · Abstract (English)

In recent times, more and more people are posting about their mental states across various social media platforms. Leveraging this data, AI-based systems can be developed that help in assessing the mental health of individuals, such as suicide risk. This paper is a study done on suicidal risk assessments using Reddit data leveraging Base language models to identify patterns from social media posts. We have demonstrated that using smaller language models, i.e., less than 500M parameters, can also be effective in contrast to LLMs with greater than 500M parameters. We propose Su-RoBERTa, a fine-tuned RoBERTa on suicide risk prediction task that utilized both the labeled and unlabeled Reddit data and tackled class imbalance by data augmentation using GPT-2 model. Our Su-RoBERTa model attained a 69.84% weighted F1 score during the Final evaluation. This paper demonstrates the effectiveness of Base language models for the analysis of the risk factors related to mental health with an efficient computation pipeline

自杀风险预测小模型社交媒体分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。