arXiv:2506.16322cs.CL2025-06中稿 · the 10th Workshop …被引 4

为波兰语大模型安全评估构建首个人工标注基准数据集。

PL-Guard: Benchmarking Language Model Safety for Polish

  • 构建波兰语安全分类数据集并生成对抗扰动样本
  • 波兰语BERT模型在对抗测试中表现最优
  • 适合关注多语言模型安全的开发者与研究者

尽管日益重视大语言模型(LLMs)的安全性,现有评估和过滤工具仍严重偏向英语等高资源语言,多数全球语言未被充分考察。为此,我们引入首个波兰语大模型安全分类的手动标注基准数据集,并创建了用于挑战模型鲁棒性的对抗扰动变体。通过一系列实验,评估了不同规模与架构的基于LLM和分类器的模型。具体包括:微调Llama-Guard-3-8B、基于HerBERT的分类器(波兰语BERT变体)以及波兰语适配的PLLuM(Llama-8B)。使用不同组合的标注数据训练模型,并与公开可用的防护模型对比性能。结果表明,HerBERT-based分类器在整体表现上最优,尤其在对抗条件下表现突出。

原文摘要 · Abstract (English)

Despite increasing efforts to ensure the safety of large language models (LLMs), most existing safety assessments and moderation tools remain heavily biased toward English and other high-resource languages, leaving majority of global languages underexamined. To address this gap, we introduce a manually annotated benchmark dataset for language model safety classification in Polish. We also create adversarially perturbed variants of these samples designed to challenge model robustness. We conduct a series of experiments to evaluate LLM-based and classifier-based models of varying sizes and architectures. Specifically, we fine-tune three models: Llama-Guard-3-8B, a HerBERT-based classifier (a Polish BERT derivative), and PLLuM, a Polish-adapted Llama-8B model. We train these models using different combinations of annotated data and evaluate their performance, comparing it against publicly available guard models. Results demonstrate that the HerBERT-based classifier achieves the highest overall performance, particularly under adversarial conditions.

模型安全波兰语多语言评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。