arXiv:2603.08275cs.CLcs.AI2026-03

让大模型生成更尊重原住民文化的内容,关键在于把文化知识嵌入安全判断。

AdaCultureSafe: Adaptive Cultural Safety Grounded by Cultural Knowledge in Large Language Models

  • 构建文化知识与安全标注配对数据集,解决评估难题
  • 发现大模型的文化安全与知识能力无显著相关性
  • 提出基于知识的响应生成方法,显著提升文化安全性

随着大语言模型(LLMs)广泛应用,尊重原住民文化已成为模型具备文化安全性和负责任全球应用的关键。现有研究分别关注文化安全与文化知识,忽视前者应以后者为基础,严重制约了模型生成特定文化尊重性回应的能力。为此,本文提出联合建模文化安全与知识的方法。首先,构建包含4.8K细粒度文化描述及48K人工验证的安全-知识导向查询的AdaCultureSafe数据集,通过权威文化知识整理、LLM自动提问生成与人工严格校验实现。在该数据集上评估三类主流大模型的文化安全与知识能力,发现二者无显著相关性。进一步分析模型内部神经元激活机制,揭示其源于预训练与后对齐目标的差异。最终提出一种知识驱动的方法,通过强制将文化知识融入生成过程,显著提升文化安全性。

原文摘要 · Abstract (English)

With the widespread adoption of Large Language Models (LLMs), respecting indigenous cultures becomes essential for models' culturally safety and responsible global applications. Existing studies separately consider cultural safety and cultural knowledge and neglect that the former should be grounded by the latter. This severely prevents LLMs from yielding culture-specific respectful responses. Consequently, adaptive cultural safety remains a formidable task. In this work, we propose to jointly model cultural safety and knowledge. First and foremost, cultural-safety and knowledge-paired data serve as the key prerequisite to conduct this research. However, the cultural diversity across regions and the subtlety of cultural differences pose significant challenges to the creation of such paired evaluation data. To address this issue, we propose a novel framework that integrates authoritative cultural knowledge descriptions curation, LLM-automated query generation, and heavy manual verification. Accordingly, we obtain a dataset named AdaCultureSafe containing 4.8K manually decomposed fine-grained cultural descriptions and the corresponding 48K manually verified safety- and knowledge-oriented queries. Upon the constructed dataset, we evaluate three families of popular LLMs on their cultural safety and knowledge proficiency, via which we make a critical discovery: no significant correlation exists between their cultural safety and knowledge proficiency. We then delve into the utility-related neuron activations within LLMs to investigate the potential cause of the absence of correlation, which can be attributed to the difference of the objectives of pre-training and post-alignment. We finally present a knowledge-grounded method, which significantly enhances cultural safety by enforcing the integration of knowledge into the LLM response generation process.

文化安全大模型知识注入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。