arXiv:2606.09178cs.CLcs.AI2026-06中稿 · ICML

跨东亚和东南亚文化适配的评测方法,让大模型安全测试更真实。

Culturally-Adapted Red-Teaming Across East and Southeast Asian Contexts: A Methodological and Comparative Analysis

  • 用文化适配替代直接翻译,构建四国语言对比数据集。
  • 文化适配提示使攻击成功率平均提升9.3个百分点。
  • 适合关注多语言模型安全与文化差异的研究者。

大型语言模型(LLM)的多语言安全评估长期依赖将英文基准直接翻译到目标语言,该方法仅转换语言表层形式,未能反映威胁场景、社会规范与法律框架中的文化内涵。本文通过1:1种子匹配,为韩语(KO)、日语(JA)、泰语(TH)和高棉语(KM)构建了成对的直接翻译(DT)与文化适配(CA)数据集,并在四个开源LLM上比较攻击成功率(ASR)与文化真实性得分。所有16个语言×模型组合中,CA提示均带来正向ΔASR(均值+9.3个百分点),而基于DT的评估在48个类别×语言组合中低估风险达44次。语言层面分析显示威胁形式分布存在异质性。文化真实性分析进一步表明,DT的文化深度(C3)得分始终低于1.0(均值0.17),而CA得分最高达2.51,说明直接翻译生成的内容与真实多文化场景显著偏离。研究证实,仅依赖语言翻译无法实现有效的多语言大模型安全评估,必须结合本地文化语境。

原文摘要 · Abstract (English)

Multilingual safety evaluation of large language models (LLMs) has predominantly relied on direct translation (DT) of English benchmarks into target languages - an approach that converts surface-level linguistic form while failing to reflect the cultural context embedded in threat scenarios, social norms, and legal frameworks. We construct paired DT and culturally-adapted (CA) datasets via 1:1 seed matching for four languages - Korean (KO), Japanese (JA), Thai (TH), and Khmer (KM) - and compare Attack Success Rate (ASR) and Cultural Realism scores across four open-source LLM. CA prompts yield Delta-ASR > 0 across all 16 language x model combinations (mean +9.3 pp), and DT-based evaluation underestimates risk in 44 of 48 category x language combinations. Language-level analysis reveals that the distribution of threat forms is heterogeneous across languages. Cultural Realism analysis further shows that DT Cultural Depth (C3) scores remain consistently below 1.0 out of 3.0 across all four languages (mean 0.17), whereas CA scores reach up to 2.51, indicating that direct translation produces inputs systematically divergent from those encountered in real-world multicultural settings. These findings demonstrate that adapting benchmarks to language-specific cultural contexts - rather than relying on linguistic translation alone - is necessary for valid multilingual LLM safety evaluation.

大模型安全文化适配多语言评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。