首个中文反仇恨话语数据集,用大模型裁判生成对抗性回应
PANDA -- Paired Anti-hate Narratives Dataset from Asia: Using an LLM-as-a-Judge to Create the First Chinese Counterspeech Dataset
- 用大模型当裁判,结合模拟退火生成中文反仇恨话语
- 构建首个面向东亚语言的反仇恨话语语料库,含文化特定攻击对象
- 适合研究中文社会媒体安全、反仇恨技术及非欧美语言的学者
尽管现代标准汉语在全球广泛使用,但针对中文的反仇恨话语(Counterspeech, CS)资源几乎空白。为填补东亚反仇恨话语研究的缺口,本文提出一种新方法,通过大模型作为裁判(LLM-as-a-Judge)、模拟退火、零样本中文生成和轮盘算法,生成针对中国大陆仇恨言论的反仇恨话语。随后进行人工验证以确保质量与语境相关性。该方法不仅揭示了中文中特定被污名化群体及程序标记为仇恨言论的语言标志模式,还提供了首个基于东亚语言的反仇恨话语语料库。分析表明,当前缺乏开源且标注完善的中文仇恨言论数据,且大模型作为裁判在中文情境下存在局限。本语料库为未来反仇恨话语生成与评估研究提供了关键资源。
原文摘要 · Abstract (English)
Despite the global prevalence of Modern Standard Chinese language, counterspeech (CS) resources for Chinese remain virtually nonexistent. To address this gap in East Asian counterspeech research we introduce the a corpus of Modern Standard Mandarin counterspeech that focuses on combating hate speech in Mainland China. This paper proposes a novel approach of generating CS by using an LLM-as-a-Judge, simulated annealing, LLMs zero-shot CN generation and a round-robin algorithm. This is followed by manual verification for quality and contextual relevance. This paper details the methodology for creating effective counterspeech in Chinese and other non-Eurocentric languages, including unique cultural patterns of which groups are maligned and linguistic patterns in what kinds of discourse markers are programmatically marked as hate speech (HS). Analysis of the generated corpora, we provide strong evidence for the lack of open-source, properly labeled Chinese hate speech data and the limitations of using an LLM-as-Judge to score possible answers in Chinese. Moreover, the present corpus serves as the first East Asian language based CS corpus and provides an essential resource for future research on counterspeech generation and evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。