首个阿尔巴尼亚语大模型安全评估数据集,填补低资源语言安全评测空白。
AlbanianLLMSafety: A Safety Evaluation Dataset for Large Language Models in Albanian
- 构建首个阿尔巴尼亚语大模型安全评测数据集,覆盖11类风险内容。
- 包含2951条提示,平均每类268条,涵盖自残、暴力、种族歧视等。
- 适用于阿尔巴尼亚语模型安全训练与防护系统开发,助力包容性AI建设。
大语言模型(LLMs)的安全评估长期聚焦高资源语言,忽视了低资源语言。本文提出 AlbanianLLMSafety,首个公开可用的阿尔巴尼亚语大模型安全评估数据集。阿尔巴尼亚语是拥有约750万使用者的独立语言,分布在阿尔巴尼亚、科索沃、北马其顿及海外侨民中。该数据集包含2951个提示,覆盖11个安全类别,包括自残、暴力、种族仇恨、儿童虐待和极端化等内容,平均每个类别268条。每条提示均提供阿尔巴尼亚语原文、英文参考翻译及详细类别标签。该资源弥补了低资源语言安全评估基础设施的重大空白,为开发更安全、更具包容性的大模型提供了关键基准。数据集将根据请求提供,用于支持阿尔巴尼亚语社区的大模型安全评估、微调、红队测试和防护机制开发。
原文摘要 · Abstract (English)
Safety evaluation of Large Language Models (LLMs) has largely focused on high-resource languages, leaving low-resource languages critically underserved. We present AlbanianLLMSafety, the first publicly available safety evaluation dataset for LLMs in Albanian, a linguistically distinct low-resource language with approximately 7.5 million speakers across Albania, Kosovo, North Macedonia, and the diaspora. The dataset contains 2,951 prompts spanning 11 safety categories, including self-harm, violence, racist content, child exploitation, and radicalization, with an average of 268 prompts per category. Each prompt is provided in Albanian with an English reference translation and a detailed category label. This resource addresses a significant gap in safety evaluation infrastruc-ture for low-resource languages and provides an essential benchmark for developing safer, more inclusive LLMs. The dataset will be provided upon request to support safety evaluation, fine-tuning, red-teaming, and guardrail development for Albanian-speaking communities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。