arXiv:2608.17827cs.CL2026-08中稿 · as non-archival pa…

为德国公共部门构建多维度模型评估框架,揭示性能与能耗、透明度间的权衡关系。

From Global Benchmarks to Local Evaluations: Benchmarking LLMs for the German Public Sector

  • 提出MÖVE框架,从能耗、厂商透明度、德党立场认知三方面评估大模型
  • 能源消耗差异超60倍,且不随模型规模线性增长
  • 欧洲模型未在德国政党立场认知上表现更优,适合关注治理合规的机构参考

公共机构在选择适配其特定场景的大语言模型时面临持续挑战。现有基准主要反映英语及美国背景,且仅评估任务性能,适用性有限。本文展示MÖVE框架的初步成果,该框架针对德国公共部门设计,综合评估三个常被忽略的治理维度:能耗、厂商透明度以及对德国政党立场的认知。结果揭示显著权衡:模型能耗差异超过60倍,且无法由模型规模解释;信息披露程度在不同厂商间系统性差异;欧洲模型并未表现出更强的德国政党立场知识。因此,公共机构的模型选择不能仅依赖性能排名,而应结合部署场景的治理要求进行综合评估。

原文摘要 · Abstract (English)

Public institutions face a persistent challenge in selecting LLMs suited to their specific context. Existing benchmarks, however, are of limited use as they primarily reflect English-language and US-centric settings, and often only evaluate task performance. In this paper, we present first results of MÖVE, a holistic evaluation framework for the German public sector, examining three rarely considered governance dimensions: energy consumption, provider transparency, and knowledge of German-party positions. Our results reveal significant trade-offs, with no single model excelling across all dimensions: estimated energy consumption varies more than 60-fold and is not explained by model size alone, information disclosure varies systematically across providers, and European models do not exhibit stronger knowledge of German party positions. Model selection for public institutions thus cannot rely on performance rankings alone. Instead, evaluations should also reflect the governance requirements of the deployment context.

大模型评估公共政策能效分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。