arXiv:2411.12767cs.CLcs.AI2024-11中稿 · publication in the…被引 1

用半监督学习提升社交媒体自杀风险评估准确率

Suicide Risk Assessment on Social Media with Semi-Supervised Learning

  • 结合500条标注数据与1500条未标注数据,改进自训练伪标签生成机制
  • 通过多轮验证筛选高质量伪标签,使模型在不平衡数据下表现更优
  • 基于RoBERTa框架,适合心理健康监测与早期干预场景

随着社交媒体成为自杀倾向者表达与聚集的场所,自然语言处理为构建自动化自杀风险评估系统提供了新路径。然而,现有研究受限于标注数据稀缺和类别分布不均的问题。为此,本文提出一种半监督框架,融合500条标注数据与1500条未标注数据,改进自训练算法中的伪标签生成过程以应对数据不平衡问题。为保障伪标签质量,人工验证了多轮生成中未达成一致的样本。测试多种模型后,最终选定RoBERTa作为骨干网络。通过使用部分验证的伪标签数据与真实标注数据联合训练,显著提升了模型从社交媒体文本中识别自杀风险的能力。

原文摘要 · Abstract (English)

With social media communities increasingly becoming places where suicidal individuals post and congregate, natural language processing presents an exciting avenue for the development of automated suicide risk assessment systems. However, past efforts suffer from a lack of labeled data and class imbalances within the available labeled data. To accommodate this task's imperfect data landscape, we propose a semi-supervised framework that leverages labeled (n=500) and unlabeled (n=1,500) data and expands upon the self-training algorithm with a novel pseudo-label acquisition process designed to handle imbalanced datasets. To further ensure pseudo-label quality, we manually verify a subset of the pseudo-labeled data that was not predicted unanimously across multiple trials of pseudo-label generation. We test various models to serve as the backbone for this framework, ultimately deciding that RoBERTa performs the best. Ultimately, by leveraging partially validated pseudo-labeled data in addition to ground-truth labeled data, we substantially improve our model's ability to assess suicide risk from social media posts.

风险评估半监督NLP心理健康

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。