arXiv:2505.04640cs.CL2025-05被引 1

对比了摩洛哥方言毒性检测模型与主流大模型的过滤效果,发现定制模型更准。

A Comparative Benchmark of a Moroccan Darija Toxicity Detection Model (Typica.ai) and Major LLM-Based Moderation APIs (OpenAI, Mistral, Anthropic)

  • 用摩洛哥方言数据集对比自研模型与三大主流API
  • 自研模型在隐性攻击和讽刺识别上F1更高
  • 适合需要本地化内容审核的平台参考

本文对比评估了Typica.ai自研摩洛哥方言毒性检测模型与主流大模型内容审核API(OpenAI omni-moderation-latest、Mistral mistral-moderation-latest、Anthropic Claude claude-3-haiku-20240307)的表现。聚焦文化相关毒性内容,包括隐性侮辱、讽刺及特定文化背景下的攻击性表达,这些常被通用系统忽略。基于从OMCD_Typica.ai_Mix数据集提取的平衡测试集,报告了精确率、召回率、F1分数和准确率,揭示了低资源语言内容审核的挑战与机遇。结果表明,Typica.ai模型表现更优,凸显了文化适配模型在可靠内容审核中的重要性。

原文摘要 · Abstract (English)

This paper presents a comparative benchmark evaluating the performance of Typica.ai's custom Moroccan Darija toxicity detection model against major LLM-based moderation APIs: OpenAI (omni-moderation-latest), Mistral (mistral-moderation-latest), and Anthropic Claude (claude-3-haiku-20240307). We focus on culturally grounded toxic content, including implicit insults, sarcasm, and culturally specific aggression often overlooked by general-purpose systems. Using a balanced test set derived from the OMCD_Typica.ai_Mix dataset, we report precision, recall, F1-score, and accuracy, offering insights into challenges and opportunities for moderation in underrepresented languages. Our results highlight Typica.ai's superior performance, underlining the importance of culturally adapted models for reliable content moderation.

内容审核方言识别AI评测本地化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。