arXiv:2410.16204cs.LGcs.CL2024-10综述被引 44

系统梳理社交网络抑郁检测中的数据与方法偏见,揭示模型泛化难题。

Machine Learning Approaches for Mental Illness Detection on Social Media: A Systematic Review of Biases and Methodological Challenges

  • 基于47篇研究系统分析机器学习在社交媒体抑郁检测中的方法缺陷
  • 超90%研究依赖英文内容,80%采用非概率采样,数据代表性差
  • 仅23%考虑语言否定等语义细节,模型易误判情绪倾向

全球精神疾病发病率上升亟需早期干预新方法。社交媒体用户生成内容为精神疾病检测提供了重要数据源。本系统综述聚焦于利用社交媒体数据通过机器学习(ML)模型检测精神疾病,特别是抑郁症,全面剖析了整个机器学习生命周期中存在的重要偏见与方法学挑战。通过检索PubMed、IEEE Xplore和Google Scholar,共识别出47篇2010年后发表的相关研究。采用预测模型偏倚风险评估工具(PROBAST)评估方法学质量与偏倚风险。研究发现,显著的偏见影响了模型的可靠性与泛化能力:多数研究过度依赖推特平台(63.8%)和英语内容(超过90%),且集中于欧美地区用户;80%采用非概率抽样,代表性不足;仅有23%的研究明确处理否定等语言细微差别,这对情感分析至关重要;27.7%的研究存在不一致的超参数调优,17%的数据划分不当,可能引发过拟合;尽管74.5%的研究使用了适合不平衡数据的评估指标,仍有部分依赖准确率而未解决类别不平衡问题,可能导致结果偏差。此外,报告透明度参差不齐,常缺少关键方法细节。因此,未来研究需拓展数据来源、统一预处理流程、确保模型开发一致性、有效应对类别不平衡,并提升报告透明度,以构建更稳健、更具普适性的社交媒体抑郁症检测模型,推动全球心理健康改善。

原文摘要 · Abstract (English)

The global increase in mental illness requires innovative detection methods for early intervention. Social media provides a valuable platform to identify mental illness through user-generated content. This systematic review examines machine learning (ML) models for detecting mental illness, with a particular focus on depression, using social media data. It highlights biases and methodological challenges encountered throughout the ML lifecycle. A search of PubMed, IEEE Xplore, and Google Scholar identified 47 relevant studies published after 2010. The Prediction model Risk Of Bias ASsessment Tool (PROBAST) was utilized to assess methodological quality and risk of bias. The review reveals significant biases affecting model reliability and generalizability. A predominant reliance on Twitter (63.8%) and English-language content (over 90%) limits diversity, with most studies focused on users from the United States and Europe. Non-probability sampling (80%) limits representativeness. Only 23% explicitly addressed linguistic nuances like negations, crucial for accurate sentiment analysis. Inconsistent hyperparameter tuning (27.7%) and inadequate data partitioning (17%) risk overfitting. While 74.5% used appropriate evaluation metrics for imbalanced data, others relied on accuracy without addressing class imbalance, potentially skewing results. Reporting transparency varied, often lacking critical methodological details. These findings highlight the need to diversify data sources, standardize preprocessing, ensure consistent model development, address class imbalance, and enhance reporting transparency. By overcoming these challenges, future research can develop more robust and generalizable ML models for depression detection on social media, contributing to improved mental health outcomes globally.

抑郁症检测社交网络机器学习偏见分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。