arXiv:2601.13558cs.AIcs.CL2026-01

用聊天记录分析男男性行为者风险行为,助力精准公共卫生干预

Leveraging ChatGPT and Other NLP Methods for Identifying Risk and Protective Behaviors in MSM: Social Media and Dating apps Text Analysis

  • 结合ChatGPT、BERT等模型分析社交平台文本
  • 预测酗酒和多伴侣行为的F1达0.78,效果良好
  • 为大规模个性化健康干预提供新方法,适合公共卫生研究者

男男性行为者(MSM)相较于异性恋男性面临更高的性传播感染和有害饮酒风险。从社交媒体和约会应用中收集的文本数据可能为个性化公共健康干预提供新机遇,实现对风险与保护行为的自动识别。本研究评估了社交平台文本在预测MSM性风险行为、饮酒情况及暴露前预防(PrEP)使用中的可行性。在获得参与者同意后,我们收集了文本数据,并利用基于ChatGPT嵌入、BERT嵌入、LIWC以及基于词典的风险词方法提取特征,训练机器学习模型。结果显示,模型在预测月度暴饮行为和拥有超过五个性伴侣方面表现优异,F1分数分别为0.78;在预测普雷佩使用和重度饮酒方面表现中等,F1分数分别为0.64和0.63。研究证实,社交平台文本可为理解风险与保护行为提供宝贵洞见,并凸显大型语言模型方法在支持可扩展、个性化公共健康干预方面的潜力。

原文摘要 · Abstract (English)

Men who have sex with men (MSM) are at elevated risk for sexually transmitted infections and harmful drinking compared to heterosexual men. Text data collected from social media and dating applications may provide new opportunities for personalized public health interventions by enabling automatic identification of risk and protective behaviors. In this study, we evaluated whether text from social media and dating apps can be used to predict sexual risk behaviors, alcohol use, and pre-exposure prophylaxis (PrEP) uptake among MSM. With participant consent, we collected textual data and trained machine learning models using features derived from ChatGPT embeddings, BERT embeddings, LIWC, and a dictionary-based risk term approach. The models achieved strong performance in predicting monthly binge drinking and having more than five sexual partners, with F1 scores of 0.78, and moderate performance in predicting PrEP use and heavy drinking, with F1 scores of 0.64 and 0.63. These findings demonstrate that social media and dating app text data can provide valuable insights into risk and protective behaviors and highlight the potential of large language model-based methods to support scalable and personalized public health interventions for MSM.

公共卫生文本分析大模型应用风险预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。