arXiv:2606.01995cs.CL2026-06

评测大模型对法国各地区细微知识的掌握能力

CARTE: A Benchmark for Mapping Language Model Knowledge Across France

  • 构建2431道题的法语区域知识测评集,覆盖13个大区14个主题
  • 12B参数模型在部分地区准确率不足60%,暴露预训练数据偏差
  • 特别设计语言变体子集,适合研究多语境下模型泛化能力

我们提出CARTE(Culturally Anchored Regional-Territorial Evaluation),一个用于评估大语言模型在法国地理语境下进行细粒度推理能力的多项选择基准。现有评测多聚焦国家层面文化理解,忽视国内区域差异及近邻地区间的区分需求。CARTE通过覆盖法国13个本土大区、14个主题(包括文化、语言、人口、经济、环境、交通等)的2431道题目填补这一空白。我们还引入CARTE-LV子集,专门评估法语地区间语言差异。在少样本设置下,我们评估了27个参数量从1B到12B不等的LLM。实验发现不同区域和模型规模间存在显著性能差异,表明预训练数据存在系统性覆盖不足,且模型对国内差异的鲁棒性有限。

原文摘要 · Abstract (English)

We introduce CARTE 1 (Culturally Anchored Regional-Territorial Evaluation), a multiplechoice benchmark for evaluating the ability of large language models (LLMs) to perform fine-grained reasoning over geographically grounded and regionally differentiated knowledge within France. While prior benchmarks focus on national-level cultural understanding, they largely overlook intra-country variation and the need to distinguish between closely related regional contexts. CARTE addresses this gap by introducing 2,431 questions spanning the 13 metropolitan regions of France and covering 14 thematic domains, including culture, language, demographics, economy, environment, and mobility. We further introduce CARTE-LV, a subset targeting Linguistic Variation across French regions, enabling focused evaluation of language-related differences. We evaluate 27 LLMs ranging from 1B to 12B parameters under few-shot settings. Our experiments reveal performance disparities across regions and model scales, suggesting systematic gaps in pretraining coverage and limited robustness to intra-national variation.

知识评测区域差异语言模型法国

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。