为阿拉伯语模型打造文化敏感的内容过滤器,兼顾安全与文化适配性。
FanarGuard: A Culturally-Aware Moderation Filter for Arabic Language Models
- 构建双语过滤框架,融合安全与文化语境评估
- 在超46万条数据上训练,文化对齐效果优于人工一致性
- 首个针对阿拉伯文化规范的评测基准,适合多语言安全研究者
内容过滤器是防止语言模型对齐失败的关键保障。然而现有方法多聚焦通用安全性,忽视文化背景。本文提出 FanarGuard,一种支持阿拉伯语和英语的双语过滤器,同时评估安全性和文化契合度。我们构建了超过46.8万条来自合成及公开数据集的提示-回复对,由大模型评审团打分,涵盖无害性和文化敏感性,并据此训练两种过滤器变体。为进一步严谨评估文化对齐,我们开发首个面向阿拉伯文化语境的基准,包含1000多个规范敏感提示,其生成回应由人类评分员标注。实验表明,FanarGuard在文化对齐方面与人工标注的一致性超越了评审员间可靠性,同时在安全基准上达到顶尖过滤器水平。结果凸显将文化意识融入过滤机制的重要性,并确立 FanarGuard 作为更上下文敏感的安全防护的实际进展。
原文摘要 · Abstract (English)
Content moderation filters are a critical safeguard against alignment failures in language models. Yet most existing filters focus narrowly on general safety and overlook cultural context. In this work, we introduce FanarGuard, a bilingual moderation filter that evaluates both safety and cultural alignment in Arabic and English. We construct a dataset of over 468K prompt and response pairs, drawn from synthetic and public datasets, scored by a panel of LLM judges on harmlessness and cultural awareness, and use it to train two filter variants. To rigorously evaluate cultural alignment, we further develop the first benchmark targeting Arabic cultural contexts, comprising over 1k norm-sensitive prompts with LLM-generated responses annotated by human raters. Results show that FanarGuard achieves stronger agreement with human annotations than inter-annotator reliability, while matching the performance of state-of-the-art filters on safety benchmarks. These findings highlight the importance of integrating cultural awareness into moderation and establish FanarGuard as a practical step toward more context-sensitive safeguards.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。