首个评估大模型文化禁忌安全的基准,发现模型知而不行。
The "Knowledge-Behavior Gap" in Cultural Taboo Safety of Large Language Models

- 构建覆盖77个国家的2020+条隐性文化禁忌数据集
- 实测多模型存在'知道但不遵守'的禁忌行为差距
- 适合关注AI伦理与跨文化安全的研究者使用
文化禁忌安全对大语言模型部署至关重要,因文化不敏感输出可能引发冒犯甚至社会伤害。然而现有文化基准多聚焦文化知识或价值观偏见,忽视模型能否识别并尊重隐含在看似无害问题中的文化禁忌。由于文化禁忌具有隐性与语境依赖性,评估面临独特挑战。为此,我们提出首个公开的专用基准CulShield,覆盖77个国家和地区,包含超过2020条禁忌条目,从显性知识与隐性行为两方面评估模型。对GPT-4o-mini、Gemini-2.5-pro等先进模型的实验揭示明显“知识-行为鸿沟”:模型虽知晓禁忌,却常在交互中未能遵循。进一步表明语言上下文差异会显著影响模型的文化禁忌安全性。代码与数据已开源:https://github.com/hedyHe/CulShield。
原文摘要 · Abstract (English)
Cultural taboo safety is essential for deploying large language models (LLMs), as culturally insensitive outputs may cause offense or even social harm. However, existing cultural benchmarks primarily assess cultural knowledge or values biases, while overlooking whether LLMs can recognize and respect cultural taboos, especially when taboos are implicitly hidden in seemingly harmless questions. Besides, cultural taboos are implicit, and context-dependent, thus poss unique challenges for reliable evaluation. To address these gaps, we introduce \textbf{CulShield}, the first public benchmark dedicated to evaluating and improving the cultural taboo safety of LLMs. CulShield spans 77 countries and territories, and includes over 2,020 taboos. It evaluates models along both explicit knowledge and implicit behaviors. Experiments on several advanced LLMs (e.g., GPT-4o-mini, Gemini-2.5-pro) reveal a clear ``knowledge-behavior gap'': models often fail to apply known taboos during interaction. We further show that variations in linguistic context can significantly affect LLMs' cultural taboo safety. Code and data is accessible here: https://github.com/hedyHe/CulShield.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。