让大模型安全测试更懂文化差异,避免翻译陷阱
CAGE: A Framework for Culturally Adaptive Red-Teaming Benchmark Generation
- 用语义模具分离攻击意图与文化内容,实现跨文化适配
- 韩语版基准KoRSET比直接翻译更能发现模型漏洞
- 适合做多语言大模型安全评估的研究者和工程师
现有红队测试基准在通过直接翻译适配新语言时,无法捕捉本地文化与法律相关的社会技术漏洞,导致大模型安全评估存在关键盲区。为此,我们提出CAGE(文化自适应生成)框架,系统性地将已验证的红队提示中的对抗性意图迁移到新文化语境。CAGE的核心是语义模具——一种将提示的对抗结构与文化内容解耦的新方法,使生成的威胁更真实、本地化,而非简单越狱测试。以韩语为例,我们构建了KoRSET基准,实验证明其在揭示模型漏洞方面显著优于直接翻译基线。CAGE为跨文化、上下文感知的安全基准开发提供了可扩展方案。数据集与评估标准已公开于https://github.com/selectstar-ai/CAGE-paper。(警告:本文包含可能具有冒犯性的模型输出。)
原文摘要 · Abstract (English)
Existing red-teaming benchmarks, when adapted to new languages via direct translation, fail to capture socio-technical vulnerabilities rooted in local culture and law, creating a critical blind spot in LLM safety evaluation. To address this gap, we introduce CAGE (Culturally Adaptive Generation), a framework that systematically adapts the adversarial intent of proven red-teaming prompts to new cultural contexts. At the core of CAGE is the Semantic Mold, a novel approach that disentangles a prompt's adversarial structure from its cultural content. This approach enables the modeling of realistic, localized threats rather than testing for simple jailbreaks. As a representative example, we demonstrate our framework by creating KoRSET, a Korean benchmark, which proves more effective at revealing vulnerabilities than direct translation baselines. CAGE offers a scalable solution for developing meaningful, context-aware safety benchmarks across diverse cultures. Our dataset and evaluation rubrics are publicly available at https://github.com/selectstar-ai/CAGE-paper. (WARNING: This paper contains model outputs that can be offensive in nature.)
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。