arXiv:2509.10886cs.CLcs.AI2025-09EMNLP被引 7

构建多语言文化问答合成框架,提升大模型跨文化理解能力

CultureSynth: A Hierarchical Taxonomy-Guided and Retrieval-Augmented Framework for Cultural Question-Answer Synthesis

  • 基于分层多语言文化分类体系与检索增强生成技术
  • 生成1.9万条跨7语言的问答数据,验证模型需30亿参数才具备基本文化能力
  • 揭示模型架构差异与地理偏差,适合开发全球化AI系统的研究者使用

文化素养指在多元文化环境中理解并适应的能力,对全球环境中的大语言模型日益重要。尽管已有多个文化评估基准,但当前评估存在分类体系零散、领域局限及高度依赖人工标注的问题。为此,我们提出CultureSynth框架,包含(1)涵盖12个主类和130个次级主题的全面分层多语言文化分类体系;(2)基于检索增强生成(RAG)的方法,利用事实知识合成文化相关问答对。CultureSynth-7合成基准包含19,360条数据,其中4,149条经人工验证,覆盖7种语言。对14种不同规模主流LLM的评估显示,性能呈现明显分层,以ChatGPT-4o-Latest和Qwen2.5-72B-Instruct领先。结果表明,实现基本文化素养需至少30亿参数,模型在知识处理上表现出不同架构偏见,且存在显著地域差异。我们认为CultureSynth为构建具有文化意识的AI系统提供了可扩展框架,同时减少了对人工标注的依赖。

原文摘要 · Abstract (English)

Cultural competence, defined as the ability to understand and adapt to multicultural contexts, is increasingly vital for large language models (LLMs) in global environments. While several cultural benchmarks exist to assess LLMs' cultural competence, current evaluations suffer from fragmented taxonomies, domain specificity, and heavy reliance on manual data annotation. To address these limitations, we introduce CultureSynth, a novel framework comprising (1) a comprehensive hierarchical multilingual cultural taxonomy covering 12 primary and 130 secondary topics, and (2) a Retrieval-Augmented Generation (RAG)-based methodology leveraging factual knowledge to synthesize culturally relevant question-answer pairs. The CultureSynth-7 synthetic benchmark contains 19,360 entries and 4,149 manually verified entries across 7 languages. Evaluation of 14 prevalent LLMs of different sizes reveals clear performance stratification led by ChatGPT-4o-Latest and Qwen2.5-72B-Instruct. The results demonstrate that a 3B-parameter threshold is necessary for achieving basic cultural competence, models display varying architectural biases in knowledge processing, and significant geographic disparities exist across models. We believe that CultureSynth offers a scalable framework for developing culturally aware AI systems while reducing reliance on manual annotation\footnote{Benchmark is available at https://github.com/Eyr3/CultureSynth.}.

文化理解多语言RAG评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。