为德国公共部门量身打造的LLM综合评测基准,兼顾性能与治理考量。
MÖVE: A Holistic LLM Benchmark for the German Public Sector

- 构建双维度评测体系:任务性能+治理合规性,覆盖39个模型
- 无模型全优表现,大小不决定质量,结果受提示词影响明显
- 适合政府选型、政策制定者及关注合规生成的开发者使用
我们提出MÖVE(Modelle für die Öffentliche Verwaltung Evaluieren),一个面向德国公共部门的大语言模型综合评测基准。尽管大语言模型在公共行政中日益普及,但模型选择仍缺乏系统指导,现有基准多以英语和美国内容为主,仅关注任务表现。MÖVE通过两个互补维度评估39个模型:性能方面涵盖摘要、问答和主题提取;治理方面评估幻觉倾向、能耗、厂商透明度及与德国宪法价值和政党立场的一致性。我们使用10个德语数据集,包括自建的金标准与银标准数据集,反映公共管理领域。采用多指标评估策略,结合传统NLP指标、嵌入方法和大模型作为裁判。结果显示,无单一模型在所有指标上领先,顶级表现因任务而异,模型规模无法有效预测质量。我们还对基准自身进行评估,分析统计精度、大模型裁判可靠性、私有数据集对排名的影响、提示词敏感性及能耗估算有效性。MÖVE为持续演进的动态基准,结果公开于https://moeve.bundesdruckerei.de/。
原文摘要 · Abstract (English)
We present MÖVE (Modelle für die Öffentliche Verwaltung Evaluieren), a holistic benchmark for evaluating large language models (LLMs) in the context of the German public sector. While LLMs are increasingly adopted in public administration, model selection remains largely ad hoc, and existing benchmarks offer limited guidance: they are predominantly English-centric, US-centric in content, and focus exclusively on task performance. MÖVE addresses these gaps by evaluating 39 models across two complementary dimensions. Performance criteria cover summarization, question answering, and topic extraction. Governance criteria assess hallucination tendencies, energy consumption, provider transparency, and alignment with German constitutional values and knowledge about positions by German political parties. In total, we utilize ten German-language datasets, including gold- and silverstandard datasets that we constructed to reflect public-administration domains. We employ a multi-metric evaluation strategy combining classical NLP metrics, embedding-based methods, and LLM-as-a-judge approaches. Our results show that no single model dominates across all criteria: top performers differ between tasks, and model size alone is a poor predictor of quality. We further evaluate the benchmark itself, analyzing its statistical precision, LLM judge reliability, the impact of our private datasets on model rankings, the sensitivity of our results to prompt formulation, and the validity of our energy consumption estimates. MÖVE is designed as a living benchmark under active development; results are publicly available at https://moeve.bundesdruckerei.de/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。