arXiv:2608.09925cs.CLcs.AI2026-08中稿 · AIES 2026

为荷兰政府设计的LLM评估框架,兼顾事实性与伦理透明度。

From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch

论文配图:From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch
图 1 · 摘自论文原文
  • 构建六维评估体系,覆盖事实性、诚实度、偏见等维度。
  • 发现高质量模型伴随更高能耗与成本,且偏见不随性能变化。
  • 区分事实性与诚实度,提出面向政策制定者的可视化工具。

大型语言模型正日益应用于政府场景,但现有评估框架难以兼顾公共管理价值与非英语语境需求。本文与荷兰主要市政机构专家合作,提出「Grip on LLMs」评估框架,通过顾问委员会、用户调研及公务人员聊天机器人使用者调查,确立六个评估维度:事实性、诚实度、社会偏见、能耗、成本和训练数据透明度,并构建涵盖30多个多语言及荷兰特有模型的基准测试套件。结果表明,无一模型在所有维度表现优异,高质量模型通常伴随更高环境影响与财务成本,而偏见独立于性能;同时发现事实性与诚实度由不同特性决定,高事实性不等于高诚实度。为便于非技术受众使用,我们发布公开可访问的友好型模型概览,供工程师到决策者全链条参考。

原文摘要 · Abstract (English)

Large language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the values of public administration and the linguistic requirements of non-English contexts. We present the "Grip on LLMs" framework, a systematic evaluation suite for Dutch governmental use developed in collaboration with domain experts from a major Dutch municipal organisation. Through an advisory board process, user research, and a survey of the users of a civil-servant chatbot, we identify six evaluation dimensions (factuality, honesty, social bias, energy consumption, cost, and training data transparency) and operationalise them into a benchmark suite covering more than 30 multilingual and Dutch-specific models. Our results reveal that no single model excels across all dimensions, and that trade-offs are unavoidable: higher quality consistently comes at greater environmental impact and financial cost, while bias remains largely independent of both. We further find that factuality (whether a model answers correctly) and honesty (whether a model acknowledges what it does not know) are governed by distinct properties, with high factuality not implying high honesty. To make these findings actionable for non-technical audiences, we release a publicly accessible, user-friendly model overview designed for the full range of stakeholders involved in governmental LLM selection, from engineers to policymakers.

LLM评估政府应用荷兰语多维度基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。