评估大模型生成反仇恨言论的四种表现,发现内容冗长且难懂。
Think Like a Person Before Responding: A Multi-Faceted Evaluation of Persona-Guided LLMs for Countering Hate
- 从角色设定、可读性、情感基调和伦理安全四方面评估反仇恨内容
- 生成内容普遍冗长,适合大学水平读者,影响传播效果
- 情感引导提升共情但安全性和有效性仍存隐患,适合研究者参考
自动反叙事(CN)是缓解网络仇恨言论的有前景策略,但其情感基调、可及性和伦理风险仍受关注。本文提出一个涵盖四种维度的评估框架:角色设定、语言冗余与可读性、情感基调、伦理鲁棒性。使用 GPT-4o-Mini、Cohere CommandR-7B 及 Meta LLaMA 3.1-70B 模型,在 MT-Conan 与 HatEval 数据集上测试三种提示策略。结果表明,生成的反叙事普遍冗长,适配大学水平读者,限制了传播可达性;情感引导提示虽提升共情与可读性,但在安全性与实际效果方面仍存在隐患。
原文摘要 · Abstract (English)
Automated counter-narratives (CN) offer a promising strategy for mitigating online hate speech, yet concerns about their affective tone, accessibility, and ethical risks remain. We propose a framework for evaluating Large Language Model (LLM)-generated CNs across four dimensions: persona framing, verbosity and readability, affective tone, and ethical robustness. Using GPT-4o-Mini, Cohere's CommandR-7B, and Meta's LLaMA 3.1-70B, we assess three prompting strategies on the MT-Conan and HatEval datasets. Our findings reveal that LLM-generated CNs are often verbose and adapted for people with college-level literacy, limiting their accessibility. While emotionally guided prompts yield more empathetic and readable responses, there remain concerns surrounding safety and effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。