针对网络毒害内容,个性化回应比通用回复更有效。
Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech

- 根据对话上下文和用户历史生成定制化反驳内容
- 结合用户特征的轻量策略提升说服力与适切性
- 实验证明个性化设计需谨慎,否则反会降低效果
AI生成的反驳内容可规模化缓解网络毒性,促进更建设性对话。然而,现有方法多采用通用、统一模式,忽视对话语境与目标用户的特性。本文提出并评估多种生成情境化反驳内容的策略,使其适应具体治理场景并个性化匹配被管理用户。我们探索了融合不同形式上下文信息与微调技术的多种配置,并通过量化指标(ROUGE、BLEU、BERTScore)与预注册的混合设计众包实验进行综合评估。结果显示,结合对话上下文与用户历史的轻量策略显著提升人类对反驳内容的适切性与说服力感知;而其他部分情境化策略反而降低了质量。研究揭示了影响说服力的关键因素,为构建更个性化、高效且负责任的反驳系统提供了可操作方向,推动人机协作在在线内容治理中的发展。
原文摘要 · Abstract (English)
AI-generated counterspeech offers a scalable and effective strategy to mitigate online toxicity by promoting more constructive dialogue. Yet, existing approaches adopt a generic, one-size-fits-all paradigm, overlooking the conversational context and characteristics of the targeted users. Here, we propose and evaluate multiple strategies for generating contextualized counterspeech that is adapted to the moderation setting and personalized to the moderated user. In detail, we explore a range of configurations that integrate different forms of contextual information and fine-tuning techniques. We conduct a comprehensive evaluation combining quantitative indicators with a pre-registered, mixed-design crowdsourcing experiment. To ensure robustness, we implement algorithmic measures of counterspeech quality based on ROUGE, BLEU, and BERTScore, observing overall consistent results across metrics. Furthermore, we analyze which characteristics of both the generated counterspeech and the moderated toxic message most strongly influence perceived persuasiveness, yielding insights into how contextualized interventions can be made more effective. Our findings show that personalization can be effective, but not uniformly so. Lightweight strategies combining conversational context and user history improve perceived adequacy and persuasiveness, whereas several other contextualization strategies degrade human-perceived counterspeech quality. Taken together, these results provide actionable directions for developing more personalized, effective, and responsible counterspeech systems, ultimately advancing human-AI collaboration in online content moderation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。