构建首个阿拉伯方言安全分类基准,评估多方言下内容安全检测效果。
ArabicDialectSafety: A Dialect-Aware Benchmark for Arabic Content Safety Classification

- 构建6种阿拉伯语方言的2.5万条标注数据集,支持细粒度危害分类。
- MARBERTv2在二分类和细粒度分类中分别达到0.95和0.90的宏平均F1值。
- 发现方言表征融合比后期处理更有效,低资源马格里布方言仍存差距。
我们提出 ArabicDialectSafety,一个包含25,071条提示的阿拉伯语安全数据集,覆盖现代标准阿拉伯语、叙利亚、埃及、阿尔及利亚、巴勒斯坦和摩洛哥六种方言,每条数据附有方言标签与七类细粒度危害标签。引入双任务评估框架,用于二元安全/不安全检测与跨方言细粒度危害分类。对七种监督模型与生成模型进行基准测试,发现微调后的 MARBERTv2 表现最佳,在二分类任务中取得 0.95 的宏平均 F1 值,细粒度分类任务中达 0.90,显著优于提示引导的前沿大模型,包括阿拉伯语专用模型。分析表明,将方言信息融入表示层最有效,但低资源马格里布方言仍存在明显性能差距。进一步评估七种前沿大模型在有害方言提示下的响应生成,观察到各模型的不安全生成率均低于 5%。数据集与代码将在论文录用后公开,以支持未来方言感知的阿拉伯语安全研究。警告:本文包含仅用于研究目的的有害及潜在冒犯性内容。
原文摘要 · Abstract (English)
We present ArabicDialectSafety, a human-curated Arabic safety dataset of 25,071 prompts covering six Arabic varieties: Modern Standard Arabic, Syrian, Egyptian, Algerian, Palestinian, and Moroccan. The dataset is annotated with dialect labels and seven fine-grained harm categories. We introduce a dual-task evaluation framework for binary safe/unsafe detection and granular harm classification across dialects. Benchmarking seven supervised and generative models, we find that fine-tuned MARBERTv2 achieves the strongest performance, with Macro-F1 scores of 0.95 for binary classification and 0.90 for granular classification, substantially outperforming prompted frontier LLMs, including Arabic-specialized models. Our analyses show that dialect conditioning is most effective when integrated at the representation level, while significant performance gaps remain for low-resource Maghrebi dialects. We further evaluate seven frontier LLMs as response generators on harmful dialectal Arabic prompts and observe unsafe generation rates below 5 percent across models. We release the dataset and code upon acceptance to support future research on dialect-aware Arabic safety evaluation. Warning: This paper contains examples of harmful and potentially offensive content included solely for research purposes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。