arXiv:2505.18927cs.CLcs.AI2025-05被引 3

评测三大模型在多语言网络欺凌检测中的表现,发现需融合模型提升准确率。

Moderating Harm: Benchmarking Large Language Models for Cyberbullying Detection in YouTube Comments

  • 统一提示下对比GPT-4.1、Gemini、Claude在多语言评论中的检测能力
  • GPT-4.1综合表现最佳(F1=0.863),Gemini漏检少但误报高,Claude精准但漏检多
  • 三模型均难识别讽刺、隐晦辱骂和跨语种俚语,适合内容审核系统优化参考

随着在线平台发展,评论区骚扰问题日益严重,影响用户体验与心理健康。本研究在包含5,080条来自游戏、生活方式、美食博主及音乐频道高骚扰帖的语料库上,对OpenAI GPT-4.1、Google Gemini 1.5 Pro和Anthropic Claude 3 Opus三个主流大模型进行基准测试。数据集涵盖1,334条有害消息与3,746条无害消息,覆盖英语、阿拉伯语和印尼语,由两名评审独立标注,一致性良好(Cohen's kappa = 0.83)。采用统一提示与确定性设置,GPT-4.1取得最优综合平衡,F1得分为0.863,精确率为0.887,召回率为0.841。Gemini召回率最高(0.875),但精确率下降至0.767,因频繁误报。Claude精确率最高(0.920),假阳性率最低(0.022),但召回率降至0.720。定性分析显示,三模型均难以识别讽刺、隐晦攻击及混合语言俚语。结果强调需构建融合互补模型、结合上下文、针对低资源语言和隐性虐待进行微调的审核流程。数据集去标识版本及完整提示已公开,以促进可复现性与自动内容审核进展。

原文摘要 · Abstract (English)

As online platforms grow, comment sections increasingly host harassment that undermines user experience and well-being. This study benchmarks three leading large language models, OpenAI GPT-4.1, Google Gemini 1.5 Pro, and Anthropic Claude 3 Opus, on a corpus of 5,080 YouTube comments sampled from high-abuse threads in gaming, lifestyle, food vlog, and music channels. The dataset comprises 1,334 harmful and 3,746 non-harmful messages in English, Arabic, and Indonesian, annotated independently by two reviewers with substantial agreement (Cohen's kappa = 0.83). Using a unified prompt and deterministic settings, GPT-4.1 achieved the best overall balance with an F1 score of 0.863, precision of 0.887, and recall of 0.841. Gemini flagged the highest share of harmful posts (recall = 0.875) but its precision fell to 0.767 due to frequent false positives. Claude delivered the highest precision at 0.920 and the lowest false-positive rate of 0.022, yet its recall dropped to 0.720. Qualitative analysis showed that all three models struggle with sarcasm, coded insults, and mixed-language slang. These results underscore the need for moderation pipelines that combine complementary models, incorporate conversational context, and fine-tune for under-represented languages and implicit abuse. A de-identified version of the dataset and full prompts is publicly released to promote reproducibility and further progress in automated content moderation.

内容审核多语言大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。