对比人与AI标注,揭示网络言论中隐性歧视的识别难题
Understanding and Analyzing Inappropriately Targeting Language in Online Discourse: A Comparative Annotation Study
- 融合众包、专家与ChatGPT标注,构建多源评估框架
- 发现语境对仇恨言论判断影响大,识别出社会信念等新靶向类别
- 指出模型在理解微妙语言时存在局限,适合内容安全研究者参考
本文提出一种检测在线对话中不当靶向语言的方法,通过整合众包标注、专家标注与ChatGPT结果,聚焦Reddit英文评论线程中针对个人或群体的言论。研究采用全面的标注框架,对多样化数据集进行目标类别及具体靶向词的标注。通过对人类专家、众包标注者与ChatGPT标注结果的对比分析,揭示了各类方法在识别显性仇恨言论与更隐蔽歧视语言方面的优劣。研究发现语境因素在仇恨言论识别中起关键作用,并发现了如社会信念、身体形象等新型靶向类别。同时,论文讨论了标注中的主观性挑战及ChatGPT在理解语言细微差别上的局限。研究成果为提升自动化内容审核策略、促进网络空间安全与包容性提供了重要启示。
原文摘要 · Abstract (English)
This paper introduces a method for detecting inappropriately targeting language in online conversations by integrating crowd and expert annotations with ChatGPT. We focus on English conversation threads from Reddit, examining comments that target individuals or groups. Our approach involves a comprehensive annotation framework that labels a diverse data set for various target categories and specific target words within the conversational context. We perform a comparative analysis of annotations from human experts, crowd annotators, and ChatGPT, revealing strengths and limitations of each method in recognizing both explicit hate speech and subtler discriminatory language. Our findings highlight the significant role of contextual factors in identifying hate speech and uncover new categories of targeting, such as social belief and body image. We also address the challenges and subjective judgments involved in annotation and the limitations of ChatGPT in grasping nuanced language. This study provides insights for improving automated content moderation strategies to enhance online safety and inclusivity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。