用机器学习提前预测健康类讨论中的毒性强言,防患于未然。
Predictively Combatting Toxicity in Health-related Online Discussions through Machine Learning
- 基于协同过滤模型,预测用户在新冠话题中可能产生的毒性言论。
- 在相关指标上预测准确率超80%,可有效识别潜在冲突组合。
- 适合平台方用于预防性干预,提升健康社区对话质量。
在健康类话题中,线上讨论中的用户毒性常引发社会冲突或传播危险、非科学行为;现有应对方式多为检测、标记或删除已有毒言论,但往往对平台和用户均不利。本文提出一种预测性对策,预先判断用户在健康类在线讨论中可能产生毒性的互动。采用基于协同过滤的机器学习方法,预测Reddit上用户与子社区间关于新冠疫情的对话毒性,相关指标预测性能超过80%,从而实现冲突用户与子社区的提前规避。
原文摘要 · Abstract (English)
In health-related topics, user toxicity in online discussions frequently becomes a source of social conflict or promotion of dangerous, unscientific behaviour; common approaches for battling it include different forms of detection, flagging and/or removal of existing toxic comments, which is often counterproductive for platforms and users alike. In this work, we propose the alternative of combatting user toxicity predictively, anticipating where a user could interact toxically in health-related online discussions. Applying a Collaborative Filtering-based Machine Learning methodology, we predict the toxicity in COVID-related conversations between any user and subcommunity of Reddit, surpassing 80% predictive performance in relevant metrics, and allowing us to prevent the pairing of conflicting users and subcommunities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。