针对中文特有攻击模式,评测轻量级大模型安全性的新基准。
CSSBench: Evaluating the Safety of Lightweight LLMs against Chinese-Specific Adversarial Patterns
- 构建涵盖六类真实场景的中文对抗模式评测集
- 发现轻量模型在中文攻击下安全失效率超60%
- 适合关注中文AI安全与轻量化部署的研究者
大型语言模型在成本敏感和本地设备场景中日益普及,但安全防护主要针对英文。真实世界中的中文恶意查询常通过谐音、拼音、符号拆分等中文特有方式隐藏意图,现有英文导向基准无法有效覆盖此类威胁,尤其对轻量级模型构成更大风险。为此,我们提出中文特定安全评测基准(CSSBench),聚焦这些对抗模式,评估轻量级中文LLM的安全性。基准涵盖非法活动与合规、隐私泄露、医疗虚假信息、欺诈仇恨、成人内容及公共政治安全六大领域,包含多种任务类型。我们评估了多款主流轻量级模型,测量其过度拒绝行为以评估安全导致的性能下降。结果显示,中文特有对抗模式是轻量级模型的关键挑战。该基准为中文场景下LLM安全性提供了全面评估工具,助力实际应用中的鲁棒部署。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed in cost-sensitive and on-device scenarios, and safety guardrails have advanced mainly in English. However, real-world Chinese malicious queries typically conceal intent via homophones, pinyin, symbol-based splitting, and other Chinese-specific patterns. These Chinese-specific adversarial patterns create the safety evaluation gap that is not well captured by existing benchmarks focused on English. This gap is particularly concerning for lightweight models, which may be more vulnerable to such specific adversarial perturbations. To bridge this gap, we introduce the Chinese-Specific Safety Benchmark (CSSBench) that emphasizes these adversarial patterns and evaluates the safety of lightweight LLMs in Chinese. Our benchmark covers six domains that are common in real Chinese scenarios, including illegal activities and compliance, privacy leakage, health and medical misinformation, fraud and hate, adult content, and public and political safety, and organizes queries into multiple task types. We evaluate a set of popular lightweight LLMs and measure over-refusal behavior to assess safety-induced performance degradation. Our results show that the Chinese-specific adversarial pattern is a critical challenge for lightweight LLMs. This benchmark offers a comprehensive evaluation of LLM safety in Chinese, assisting robust deployments in practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。