arXiv:2410.01400cs.CL2024-10被引 7

首个针对六类反仇恨言论的多类型数据集,助力高效生成精准回应。

CrowdCounter: A benchmark type-specific multi-target counterspeech dataset

  • 构建六类反仇恨言论回应数据集,鼓励高质量非重复输出。
  • 使用类型控制提示提升回应相关性,但可能降低语言质量。
  • 适合研究反仇恨言论生成、平台内容治理的学者与工程师。

反仇恨言论为在不封禁用户的情况下维护言论自由提供了可行方案,但撰写有效回应对审核员和用户而言仍具挑战。因此,开发辅助生成反仇恨言论回应的工具迫在眉睫。现有数据集面临响应质量与多样性不足的问题。为此,我们提出新数据集 CrowdCounter,包含3,425对仇恨言论-反仇恨言论配对,覆盖六种回应类型(共情、幽默、质疑、警告、羞辱、反驳),为首个此类数据集。其标注平台设计促进类型特定、非冗余且高质量的回应生成。我们在四个大语言模型上评估了两种生成框架:普通提示与类型控制提示。通过相关性、多样性和质量三项指标进行评测,发现Flan-T5在普通框架中表现最佳;类型控制提示显著提升相关性,但可能降低语言质量;DialoGPT在遵循指令并准确生成类型化回应方面表现最优。

原文摘要 · Abstract (English)

Counterspeech presents a viable alternative to banning or suspending users for hate speech while upholding freedom of expression. However, writing effective counterspeech is challenging for moderators/users. Hence, developing suggestion tools for writing counterspeech is the need of the hour. One critical challenge in developing such a tool is the lack of quality and diversity of the responses in the existing datasets. Hence, we introduce a new dataset - CrowdCounter containing 3,425 hate speech-counterspeech pairs spanning six different counterspeech types (empathy, humor, questioning, warning, shaming, contradiction), which is the first of its kind. The design of our annotation platform itself encourages annotators to write type-specific, non-redundant and high-quality counterspeech. We evaluate two frameworks for generating counterspeech responses - vanilla and type-controlled prompts - across four large language models. In terms of metrics, we evaluate the responses using relevance, diversity and quality. We observe that Flan-T5 is the best model in the vanilla framework across different models. Type-specific prompts enhance the relevance of the responses, although they might reduce the language quality. DialoGPT proves to be the best at following the instructions and generating the type-specific counterspeech accurately.

反仇恨言论数据集语言模型内容安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。