arXiv:2608.12894cs.CL2026-08

评测大模型对巴伐利亚方言与地方文化的理解能力,发现模型在方言上表现较差。

BavGround: A Benchmark for Regional Cultural Grounding and Dialect Competence in Bavarian

论文配图:BavGround: A Benchmark for Regional Cultural Grounding and Dialect Competence in Bavarian
图 1 · 摘自论文原文
  • 构建多语言巴伐利亚文化评测集,涵盖8个文化领域共618个题目。
  • 模型在巴伐利亚语和来源依赖型题目上表现显著下降,显示方言理解不足。
  • 不同评估方法导致结果差异大,建议采用协议敏感的评测方式。

大型语言模型的文化评估多集中于高资源标准语言,忽视了区域性文化与方言群体。本文提出BavGround,一个用于评估英语、德语和巴伐利亚语中巴伐利亚地区文化认知与方言能力的基准测试。该基准包含每种语言8个文化领域共206道多选题,形成618个平行实例,题目覆盖广泛可及的文化知识以及来自新闻、历史文献和专业著作的源文本支持的地方性知识。我们评估了15个70亿至100亿参数的开源指令微调模型和1个闭源模型。强健的多语言模型整体表现最佳,但在巴伐利亚语题目和源文本依赖题上性能下降明显,表明模型仍难以处理方言与本地化文化知识。进一步分析显示,评估协议影响显著:原始答案字母评分、打乱字母评分、选项文本似然、生成答案解析和语义匹配等方法会产生不同的绝对得分与排名,尤其对区域适配模型影响更大。对GENBA-10B检查点的探索性分析表明,持续预训练虽提升了部分领域的答案内容似然,但方言能力改善有限。BavGround支持针对地方文化表征的精细化、协议敏感型评估。

原文摘要 · Abstract (English)

Cultural evaluation of large language models (LLMs) often focuses on high-resource standard languages, leaving regional culture and dialect communities underrepresented. We introduce BavGround, a benchmark for evaluating Bavarian regional cultural grounding and dialect competence across English, German and Bavarian. BavGround contains 206 multiple-choice source questions across eight cultural domains per language, yielding 618 multi-parallel instances, with items covering both broadly accessible cultural knowledge and source-grounded regional knowledge from journalism, historical sources, and specialist literature. We evaluate fifteen 7B-10B open-weight instruction-tuned models and one closed-model reference. Strong multilingual models perform best overall, but performance drops on Bavarian items and source-grounded questions, indicating persistent difficulty with dialectal and localized cultural knowledge. We further show that conclusions depend strongly on evaluation protocol: raw answer-letter scoring, shuffled-letter scoring, option-text likelihood, generated-answer parsing, and semantic matching can produce different absolute scores and rankings, especially for regionally adapted models. Finally, an exploratory analysis of GENBA-10B checkpoints suggests that continued pretraining improves answer-content likelihood unevenly across domains, while dialect competence remains comparatively weak. BavGround supports localized, protocol-aware evaluation of cultural representation in LLMs.

文化认知方言评测多语言基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。