针对韩文化特有风险,构建多模态安全评估基准,发现模型对本土化攻击更脆弱。
KSAFE-MM: A Multimodal Safety Benchmark via Localized Contextualization for Korean Cultural Risks

- 通过语言语境化生成韩国场景下的通用安全测试样例
- 用真实文化视觉+越狱文本组合测试本土化攻击成功率达74.2%
- 揭示安全与过度拒绝的权衡,适合关注跨文化AI安全的研究者
多模态大语言模型(MLLMs)在语言和视觉等多模态中引入安全漏洞。现有安全评估工具存在两大局限:一是以英语为中心的数据集构建,二是仅关注通用风险而缺乏本地文化关联。本文提出KSAFE-MM,一个面向韩语文化的多模态安全评估基准,涵盖通用风险与文化特异性漏洞。该基准包含两部分:KSAFE-MM-G通过语言语境化将通用安全问题转化为韩国语境下的多模态样本;KSAFE-MM-C则基于真实场景的本土化视觉查询,搭配越狱式文本查询,覆盖涉及文化视觉线索与恶意文本意图的多模态风险。在12个先进MLLM上评估显示,模型对文化相关攻击的脆弱性高于通用攻击,其中程序执行类越狱策略使攻击成功率达74.2%,远超标准查询的13.4%。此外,低攻击成功率模型往往表现出对正常查询的过度拒绝,揭示安全与过拒间的系统性权衡。研究强调亟需超越英语中心的跨文化安全评估。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) exacerbate safety risks by introducing vulnerabilities across multiple modalities, such as language and vision. Current MLLM safety evaluation tools, however, suffer from major limitations: 1) English-centric dataset construction, and 2) a focus on generic risks that are not tied to local cultural contexts. This paper introduces KSAFE-MM, a benchmark for Korean multimodal safety evaluation that covers both general safety risks and culture-specific vulnerabilities. KSAFE-MM consists of two parts, KSAFE-MM-G and KSAFE-MM-C. KSAFE-MM-G evaluates globally shared risks in Korean contexts through linguistic contextualization, which transforms generic safety queries into contextually grounded multimodal samples. KSAFE-MM-C targets culture-dependent MLLM safety vulnerabilities using localized visual queries derived from real-world contexts. It pairs these visual queries with jailbreak-style textual queries to cover multimodal safety risks involving cultural visual cues and malicious textual intent. Together, these components provide a general-to-local construction pipeline for evaluating both globally shared safety risks and culture-specific vulnerabilities. We evaluate 12 state-of-the-art MLLMs on KSAFE-MM and reveal that models exhibit greater vulnerability to culturally grounded attacks than to generic ones. Notably, jailbreaking strategies substantially amplify attack success rates, with ProgramExecution yielding up to 74.2% ASR compared to 13.4% for standard queries. Furthermore, we identify a systematic trade-off between safety and over-refusal, where models achieving low ASR tend to exhibit excessive refusal behavior on benign queries. These findings highlight the urgent need for culturally grounded safety evaluation beyond English-centric benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。