评测大模型生成反仇恨言论的逼真度,发现机器与真人写法可轻易区分。
Assessing the Human Likeness of AI-Generated Counterspeech
- 用大模型生成反仇恨内容,对比人类写作差异。
- 机器生成内容在语气、具体性上与真人有明显区别。
- 结果对改进AI反垃圾内容系统有重要参考价值。
反仇恨言论是针对仇恨或攻击性内容的针对性回应,能有效遏制负面信息传播并促进良性网络交流。以往研究提出多种自动生成反仇恨言论的方法,但评估多聚焦于相关性、表面形式等浅层语言特征。本文首次系统考察AI生成反仇恨言论的逼真度这一关键影响因素。我们实现了多种基于大语言模型的生成策略,并通过简单分类器和人工评估发现,AI生成内容与真人撰写的内容可被轻易区分。进一步分析揭示两者在语言特征、礼貌程度和具体性方面存在显著差异。本研究使用的数据集已公开,可供后续研究使用。
原文摘要 · Abstract (English)
Counterspeech is a targeted response to counteract and challenge abusive or hateful content. It effectively curbs the spread of hatred and fosters constructive online communication. Previous studies have proposed different strategies for automatically generated counterspeech. Evaluations, however, focus on relevance, surface form, and other shallow linguistic characteristics. This paper investigates the human likeness of AI-generated counterspeech, a critical factor influencing effectiveness. We implement and evaluate several LLM-based generation strategies, and discover that AI-generated and human-written counterspeech can be easily distinguished by both simple classifiers and humans. Further, we reveal differences in linguistic characteristics, politeness, and specificity. The dataset used in this study is publicly available for further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。