arXiv:2412.07338cs.HCcs.AI2024-12被引 28

让AI回复更懂语境,提升反偏见言论的说服力。

Contextualized Counterspeech: Strategies for Adaptation, Personalization, and Evaluation

  • 用上下文和用户特征定制回复,取代通用模板
  • 实测显示定制回复在说服力上显著优于现有方法
  • 强调人工评估与算法指标差异,需更精细评价体系

AI生成的反向言论可有效缓解网络暴力,但现有方法缺乏对语境和用户的适应性。本文提出多种策略,基于LLaMA2-13B模型生成针对具体情境与目标用户的个性化反向言论。通过预注册的混合设计众包实验,结合定量指标与人工评估,验证了不同配置的效果。结果表明,上下文感知的反向言论在恰当性和说服力上显著优于当前最优的通用模型,且未牺牲其他特性。研究还发现,量化指标与人工评价相关性较差,说明二者衡量维度不同,凸显了建立更细致评估方法的必要性。该研究强调了人机协作在内容审核中的关键作用。

原文摘要 · Abstract (English)

AI-generated counterspeech offers a promising and scalable strategy to curb online toxicity through direct replies that promote civil discourse. However, current counterspeech is one-size-fits-all, lacking adaptation to the moderation context and the users involved. We propose and evaluate multiple strategies for generating tailored counterspeech that is adapted to the moderation context and personalized for the moderated user. We instruct an LLaMA2-13B model to generate counterspeech, experimenting with various configurations based on different contextual information and fine-tuning strategies. We identify the configurations that generate persuasive counterspeech through a combination of quantitative indicators and human evaluations collected via a pre-registered mixed-design crowdsourcing experiment. Results show that contextualized counterspeech can significantly outperform state-of-the-art generic counterspeech in adequacy and persuasiveness, without compromising other characteristics. Our findings also reveal a poor correlation between quantitative indicators and human evaluations, suggesting that these methods assess different aspects and highlighting the need for nuanced evaluation methodologies. The effectiveness of contextualized AI-generated counterspeech and the divergence between human and algorithmic evaluations underscore the importance of increased human-AI collaboration in content moderation.

AI反制内容审核个性化生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。