用联邦学习保护社交网络心理健康数据隐私,但差分隐私会大幅降低检测效果。
FedMental: Evaluating Federated Learning for Mental Health Detection from Social Media Data

- 将用户视为客户端,在非独立同分布场景下模拟联邦学习
- 联邦学习在抑郁检测上表现接近集中训练(F1 83.16 vs 85.63)
- 差分隐私导致性能下降超50%(最低F1 27.01),因关键语言特征被破坏
社交平台文本常用于训练机器学习模型以识别高风险心理健康行为。然而,共享此类敏感数据存在隐私风险,并限制基准数据集的发展。本文系统评估了隐私保护机器学习技术能否在保障安全的同时维持性能。具体地,针对推特上的抑郁检测和Reddit上的自杀危机检测两个典型任务,应用联邦学习(FL)与差分隐私联邦学习(DP-FL)。通过将每位用户视为客户端,在非独立同分布设置下,评估不同客户端比例、聚合策略及隐私预算下的表现。结果表明,联邦学习在抑郁识别任务中表现接近集中训练(集中式F1=85.63;最优FL模型F1=83.16);但差分隐私联邦学习即使在低噪声水平(ε=50)下也出现显著性能-隐私权衡(最高F1下降至27.01),原因在于高度信息量但稀疏的心理健康语言标记(如情绪词、健康话题)被严重扭曲。该研究实证揭示了当前隐私保护技术在心理健康推断任务中的潜力与局限。
原文摘要 · Abstract (English)
Social media text data are often used to train Machine Learning (ML) models to identify users exhibiting high-risk mental health behaviors. However, sharing this sensitive data poses privacy risks and limits the growth of benchmark datasets. We comprehensively evaluate whether privacy-preserving ML techniques can enable safer data sharing while preserving performance. Specifically, we apply federated learning (FL) and Differentially Private FL for two widely-studied mental health prediction tasks: depression detection on X (Twitter) and suicide crisis detection on Reddit. We simulate realistic data-sharing scenarios by treating each user as a client in a non-IID setting, evaluating across different client fractions, aggregation strategies, and privacy budgets. While FL achieves comparable performance to centralized training (centralized F1 = 85.63; best FL model F1 = 83.16) on depression identification, we find that Differentially Private FL has a large performance-privacy trade-off (up to F1 = 27.01 drop) even with low levels of noise (epsilon = 50). This is due to the distortion of highly informative yet sparse mental health linguistic markers related to mental health, like health topics and emotion words. This research empirically demonstrates the potential and limitations of current privacy preservation techniques for mental health inference tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。