arXiv:2601.20100cs.HCcs.AI2026-01

用AI聊天机器人干预网络毒言,尝试重塑用户行为

Taming Toxic Talk: Using chatbots to intervene with users posting toxic comments

  • 设计AI对话引导曾发毒语的用户反思
  • 893人参与实验,但一个月内毒言行为无显著减少
  • 适合研究在线治理与人机互动的学者和平台方

生成式AI聊天机器人在实验室中已证明能有效改变人的信念与态度。本文探讨了使用生成式AI聊天机器人对发布有毒内容的用户进行康复性对话的实际效果。毒性行为(如人身攻击或暴力威胁)在在线社区中普遍存在,现有应对策略多为惩罚性措施,如删帖或封号。由于与攻击性用户互动的心理成本高,康复性方法很少被采用。本研究与七个大型Reddit社区合作,开展大规模实地实验(N=893),邀请近期发布毒言的用户参与与AI聊天机器人的对话。定性分析显示,许多参与者真诚回应,甚至表达悔意或改变意愿。然而,与对照组相比,实验组在接下来的一个月内并未表现出显著的毒性行为减少。本文讨论了结果可能的原因,以及由此带来的理论与实践启示。

原文摘要 · Abstract (English)

Generative AI chatbots have proven surprisingly effective at persuading people to change their beliefs and attitudes in lab settings. However, the practical implications of these findings are not yet clear. In this work, we explore the impact of rehabilitative conversations with generative AI chatbots on users who share toxic content online. Toxic behaviors -- like insults or threats of violence, are widespread in online communities. Strategies to deal with toxic behavior are typically punitive, such as removing content or banning users. Rehabilitative approaches are rarely attempted, in part due to the emotional and psychological cost of engaging with aggressive users. In collaboration with seven large Reddit communities, we conducted a large-scale field experiment (N=893) to invite people who had recently posted toxic content to participate in conversations with AI chatbots. A qualitative analysis of the conversations shows that many participants engaged in good faith and even expressed remorse or a desire to change. However, we did not observe a significant change in toxic behavior in the following month compared to a control group. We discuss possible explanations for our findings, as well as theoretical and practical implications based on our results.

AI干预在线治理毒言行为人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。