arXiv:2605.26947cs.CL2026-05中稿 · the SIGUL2026 Work…

首个哈萨克语安全评估提示数据集,填补低资源语言安全评测空白。

KZ-SafetyPrompts: A Kazakh Safety Evaluation Prompt Dataset for Large Language Models

  • 构建11类哈萨克语安全提示,覆盖自残、暴力等风险场景。
  • 5717条原生哈萨克语提示,GPT-4o拒绝率为5.5%~53.8%不等。
  • 专为低资源语言设计,适合多语言安全评测研究者使用。

哈萨克语在大语言模型安全行为评估资源中严重不足。本文提出KZ-SafetyPrompts,一个包含5717条原生哈萨克语(西里尔字母)提示的评估数据集,涵盖自残、暴力、儿童剥削、性内容、种族歧视、极端化及管制品或非法活动等11个风险类别。提示模拟真实用户查询,常以青少年或儿童口吻表达意图,无具体操作指令。提供英文翻译以支持跨语言分析,并记录编写规范、标注流程(含边缘案例判定规则)及质控步骤(模式标准化、完整性检查、去重)。类别与主流安全分类体系对齐,便于集成现有评估流程。基线测试显示GPT-4o总体拒绝率为28.2%,各类别间波动于5.5%至53.8%,表明哈萨克语提示揭示了英语单语评估无法捕捉的特定类别安全漏洞。

原文摘要 · Abstract (English)

Kazakh is underrepresented in resources for evaluating the safety behavior of large language models. We present KZ-SafetyPrompts, a Kazakh prompt dataset for safety evaluation across eleven categories covering common risk areas such as self-harm, violence, child exploitation, sexual content, racist content, radicalization, and regulated goods or illegal activities. The dataset contains 5,717 prompts written natively in Kazakh (Cyrillic), organized by category, with English translations for cross-lingual analysis. Prompts resemble realistic user queries, often in a teen or child style, and are phrased as intent prompts without procedural instructions. We document the writing protocol, labeling procedures (including borderline-case decision rules), and quality-control steps (schema standardization, completeness checks, and deduplication). We also align the categories with widely used safety taxonomies to support integration with existing evaluation pipelines. Baseline results with GPT-4o show an overall refusal rate of 28.2%, varying from 5.5% to 53.8% across categories, indicating that Kazakh prompts expose category-specific safety gaps not captured by English-only evaluation.

安全评测低资源语言提示数据集多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。