安全防护反而削弱大模型反仇恨言论的论证力,针对性攻击更能提升效果。
Is Safer Better? The Impact of Guardrails on the Argumentative Strength of LLMs in Hate Speech Countering
- 测试安全机制对生成反驳内容的影响,发现其可能抑制论证深度。
- 针对性攻击隐含偏见和仇恨部分,生成内容质量更高。
- 适合关注AI伦理与社会影响、反仇恨策略的研究者参考。
自动生成反驳性言论作为遏制仇恨言论的策略日益受到自然语言生成领域关注,但现有生成结果常缺乏专家级反驳所具备的论证深度。本文从两个方面提升反驳生成质量:首先,探讨大模型在有用性与无害性之间的权衡,验证安全防护机制是否妨碍生成质量;其次,评估针对仇恨言论特定组成部分(如隐含负面刻板印象或仇恨内容)进行攻击是否能带来更有效的论证策略。通过大规模人工与自动评估,研究发现安全防护机制可能对本应促进积极社会互动的任务产生负面影响;同时,针对具体组件的攻击策略显著提升了生成内容的质量。
原文摘要 · Abstract (English)
The potential effectiveness of counterspeech as a hate speech mitigation strategy is attracting increasing interest in the NLG research community, particularly towards the task of automatically producing it. However, automatically generated responses often lack the argumentative richness which characterises expert-produced counterspeech. In this work, we focus on two aspects of counterspeech generation to produce more cogent responses. First, by investigating the tension between helpfulness and harmlessness of LLMs, we test whether the presence of safety guardrails hinders the quality of the generations. Secondly, we assess whether attacking a specific component of the hate speech results in a more effective argumentative strategy to fight online hate. By conducting an extensive human and automatic evaluation, we show how the presence of safety guardrails can be detrimental also to a task that inherently aims at fostering positive social interactions. Moreover, our results show that attacking a specific component of the hate speech, and in particular its implicit negative stereotype and its hateful parts, leads to higher-quality generations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。