arXiv:2608.01291cs.CL2026-08

构建首个阿拉伯方言安全分类基准,评估多方言下内容安全检测效果。

ArabicDialectSafety: A Dialect-Aware Benchmark for Arabic Content Safety Classification

论文配图:ArabicDialectSafety: A Dialect-Aware Benchmark for Arabic Content Safety Classification
图 1 · 摘自论文原文
  • 构建6种阿拉伯语方言的2.5万条标注数据集,支持细粒度危害分类。
  • MARBERTv2在二分类和细粒度分类中分别达到0.95和0.90的宏平均F1值。
  • 发现方言表征融合比后期处理更有效,低资源马格里布方言仍存差距。

我们提出 ArabicDialectSafety,一个包含25,071条提示的阿拉伯语安全数据集,覆盖现代标准阿拉伯语、叙利亚、埃及、阿尔及利亚、巴勒斯坦和摩洛哥六种方言,每条数据附有方言标签与七类细粒度危害标签。引入双任务评估框架,用于二元安全/不安全检测与跨方言细粒度危害分类。对七种监督模型与生成模型进行基准测试,发现微调后的 MARBERTv2 表现最佳,在二分类任务中取得 0.95 的宏平均 F1 值,细粒度分类任务中达 0.90,显著优于提示引导的前沿大模型,包括阿拉伯语专用模型。分析表明,将方言信息融入表示层最有效,但低资源马格里布方言仍存在明显性能差距。进一步评估七种前沿大模型在有害方言提示下的响应生成,观察到各模型的不安全生成率均低于 5%。数据集与代码将在论文录用后公开,以支持未来方言感知的阿拉伯语安全研究。警告:本文包含仅用于研究目的的有害及潜在冒犯性内容。

原文摘要 · Abstract (English)

We present ArabicDialectSafety, a human-curated Arabic safety dataset of 25,071 prompts covering six Arabic varieties: Modern Standard Arabic, Syrian, Egyptian, Algerian, Palestinian, and Moroccan. The dataset is annotated with dialect labels and seven fine-grained harm categories. We introduce a dual-task evaluation framework for binary safe/unsafe detection and granular harm classification across dialects. Benchmarking seven supervised and generative models, we find that fine-tuned MARBERTv2 achieves the strongest performance, with Macro-F1 scores of 0.95 for binary classification and 0.90 for granular classification, substantially outperforming prompted frontier LLMs, including Arabic-specialized models. Our analyses show that dialect conditioning is most effective when integrated at the representation level, while significant performance gaps remain for low-resource Maghrebi dialects. We further evaluate seven frontier LLMs as response generators on harmful dialectal Arabic prompts and observe unsafe generation rates below 5 percent across models. We release the dataset and code upon acceptance to support future research on dialect-aware Arabic safety evaluation. Warning: This paper contains examples of harmful and potentially offensive content included solely for research purposes.

内容安全阿拉伯语多语言方言识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。