arXiv:2605.29667cs.CL2026-05被引 1

针对中文场景设计的高风险模型安全评测基准,发现英文安全系统在中文下失效。

Beyond English and Evasion: A Human-Annotated Multi-Domain Benchmark for High-Stakes LLM Safety Evaluation in Chinese

  • 构建1897条中文对抗性提示,覆盖自残、毒品、欺诈等四大高风险领域
  • 1544条标注数据含响应类型、混淆手法分类与风险等级,支持精细化评估
  • 揭示语言文化差异对安全系统的影响,适合研究中文LLM安全与对齐的团队

当大型语言模型部署于中文环境时,一个令人担忧的现象浮现:在英文中表现良好的安全系统在中文中失效。这些系统难以跨越语言与文化边界,无法抵御利用拼音转写、字符拆解、网络用语和语气模糊等中文特有规避技术的攻击。为此,我们提出ChiSafe-PAS(中文安全试点标注集),一个包含1897条跨领域中文对抗性提示的人工标注基准,涵盖自伤与暴力、毒品与非法交易、诈骗及讽刺四类高风险场景。其中1544条具备完整黄金标准标注:三类响应标签(拒绝、安全引导、回应)、九类混淆手法分类、风险等级评分及标注者推理。本文详述数据集设计、标注流程与混淆分类体系。核心目标是为研究社区提供高质量、文化语境贴合的模型安全对齐评测资源。同时,本工作也触及领域内三大张力:训练与评测数据边界的模糊化、真实风险驱动的领域覆盖需求,以及规模无法替代文化专长的局限。

原文摘要 · Abstract (English)

When Large Language Models (LLMs) are deployed in Chinese-language settings, a troubling pattern emerges: safety systems that work well in English break down. These systems struggle to cross linguistic and cultural bound-aries, leaving models exposed to adversarial prompts that exploit Chinese-specific evasion techniques, including Pinyin romanization, character decomposition, internet slang, and hedging tone. To address this gap, we introduce ChiSafe-PAS (Chinese Safety Pilot Annotation Set), a human-annotated benchmark of 1,897 adversarial Chinese prompts spanning four high-stakes domains: self-harm and violence, drug and illicit trade, fraud, and satire. Of these, 1,544 entries carry complete gold-standard annotations: a 3-class response label (REFUSE, SAFE-REDIRECT, RESPOND), a nine-category obfuscation taxonomy, a risk-level rating, and annotator rationale. We describe the dataset design, annotation process, and obfuscation taxonomy in detail. Our primary goal is practical: to give the research community a high-quality, culturally grounded resource for benchmarking LLM safety alignment. In doing so, we engage three broader tensions in the field: the blurring boundary between training and evaluation data, the need for domain coverage grounded in real-world risk, and the limits of scale as a substitute for cultural expertise.

模型安全中文LLM对抗样本评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。