通过多属性条件生成更精准的反仇恨言论,提升回应有效性。
Counterspeech the ultimate shield! Multi-Conditioned Counterspeech Generation through Attributed Prefix Learning
- 分层前缀学习融合多种属性,动态优化生成过程
- 相比基线模型,意图符合度提升38%,文本质量显著改善
- 适合研究内容安全、社交媒体治理的学者与开发者
反仇恨言论已被证明是应对网络仇恨言论的有效工具。以往研究仅基于单一意图生成反言论(单属性条件),而同时考虑多个属性的综合方法可产生更细致且有效的回应。本文提出HiPPrO框架——一种两阶段方法,第一阶段通过层次化前缀学习在生成过程中逐步优化特定属性嵌入空间;第二阶段引入参考和无奖励偏好优化,生成更具建设性的反言论。我们扩展了IntentCONANv2数据集,由五名标注者为全部13,973条反言论标注情绪标签。HiPPrO利用层级前缀优化有效整合意图与情绪双重属性。大量评估显示,相较多个基线模型,其意图符合度提升约38%,Rouge-1、Rouge-2、Rouge-L分别提升约3%、2%、3%。人工评估进一步证实本方法在相关性与适切性上的优势。该工作揭示了多属性条件化在提升反言论生成系统效能方面的潜力。代码已开源于Github,数据集在Hugging Face开放。
原文摘要 · Abstract (English)
Counterspeech has proven to be a powerful tool to combat hate speech online. Previous studies have focused on generating counterspeech conditioned only on specific intents (single attributed). However, a holistic approach considering multiple attributes simultaneously can yield more nuanced and effective responses. Here, we introduce HiPPrO, Hierarchical Prefix learning with Preference Optimization, a novel two-stage framework that utilizes the effectiveness of attribute-specific prefix embedding spaces hierarchically optimized during the counterspeech generation process in the first phase. Thereafter, we incorporate both reference and reward-free preference optimization to generate more constructive counterspeech. Furthermore, we extend IntentCONANv2 by annotating all 13,973 counterspeech instances with emotion labels by five annotators. HiPPrO leverages hierarchical prefix optimization to integrate these dual attributes effectively. An extensive evaluation demonstrates that HiPPrO achieves a ~38 % improvement in intent conformity and a ~3 %, ~2 %, ~3 % improvement in Rouge-1, Rouge-2, and Rouge-L, respectively, compared to several baseline models. Human evaluations further substantiate the superiority of our approach, highlighting the enhanced relevance and appropriateness of the generated counterspeech. This work underscores the potential of multi-attribute conditioning in advancing the efficacy of counterspeech generation systems. Our code is available on Github and dataset is open-sourced on Hugging-face.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。