用大模型辅助非母语者跨文化仇恨言论审核,准确率超GPT-4o且人力减少83.6%。
LLM-C3MOD: A Human-LLM Collaborative System for Cross-Cultural Hate Speech Moderation
- 三步协同流程:增强文化注释、大模型初筛、人工介入分歧项
- 在韩语数据集上达78%准确率,超越GPT-4o的71%基准
- 适合需跨语言审核的平台与缺乏本地化人力的团队
内容审核是全球性挑战,但主流平台优先高资源语言,导致低资源语言缺乏本地审核员。由于有效审核依赖上下文理解,非母语者因文化认知不足,易误判。用户研究发现,非母语审核员在解读文化特定知识、情感及网络文化时存在困难。为此,我们提出LLM-C3MOD,一种人机协同系统,包含三步:(1) 基于RAG的文化背景增强标注;(2) 大模型初步审核;(3) 对模型无共识的案例进行精准人工干预。在韩语仇恨言论数据集上,由印尼和德国参与者测试,系统达到78%准确率(优于GPT-4o的71%),同时将人工工作量降低83.6%。值得注意的是,人类在处理细微语义时表现更优。结果表明,经大模型支持后,非母语审核员可有效参与跨文化审核。
原文摘要 · Abstract (English)
Content moderation is a global challenge, yet major tech platforms prioritize high-resource languages, leaving low-resource languages with scarce native moderators. Since effective moderation depends on understanding contextual cues, this imbalance increases the risk of improper moderation due to non-native moderators' limited cultural understanding. Through a user study, we identify that non-native moderators struggle with interpreting culturally-specific knowledge, sentiment, and internet culture in the hate speech moderation. To assist them, we present LLM-C3MOD, a human-LLM collaborative pipeline with three steps: (1) RAG-enhanced cultural context annotations; (2) initial LLM-based moderation; and (3) targeted human moderation for cases lacking LLM consensus. Evaluated on a Korean hate speech dataset with Indonesian and German participants, our system achieves 78% accuracy (surpassing GPT-4o's 71% baseline), while reducing human workload by 83.6%. Notably, human moderators excel at nuanced contents where LLMs struggle. Our findings suggest that non-native moderators, when properly supported by LLMs, can effectively contribute to cross-cultural hate speech moderation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。