arXiv:2601.20617cs.CYcs.AI2026-01被引 3

现有评测基准无法满足公共部门对LLM代理的法律与流程要求。

Agent Benchmarks Fail Public Sector Requirements

  • 从公共管理理论出发,定义评测基准需具备流程化、真实性和领域特异性。
  • 分析1300多篇基准论文后发现:无一完全符合四大核心标准。
  • 呼吁研究者开发适配公共部门的评测体系,官员评估代理时应参考这些标准。

在公共部门部署基于大语言模型的智能体(LLM代理)需确保其满足严格的法律、程序与结构要求。从业者和研究者常依赖评测基准进行评估,但现有基准是否足以反映公共部门需求尚不明确。本文基于公共行政学文献,提出四项核心标准:基准必须是流程导向的、现实的、专属于公共部门的,并报告能体现公共部门独特要求的指标。通过专家验证的LLM辅助管道,我们分析了超过1,300篇基准论文。结果表明,没有任何一个基准同时满足全部四项标准。研究呼吁研究人员开发更贴合公共部门需求的评测体系,同时建议公共部门官员在评估代理应用时采用这些标准。

原文摘要 · Abstract (English)

Deploying Large Language Model-based agents (LLM agents) in the public sector requires assuring that they meet the stringent legal, procedural, and structural requirements of public-sector institutions. Practitioners and researchers often turn to benchmarks for such assessments. However, it remains unclear what criteria benchmarks must meet to ensure they adequately reflect public-sector requirements, or how many existing benchmarks do so. In this paper, we first define such criteria based on a first-principles survey of public administration literature: benchmarks must be \emph{process-based}, \emph{realistic}, \emph{public-sector-specific} and report \emph{metrics} that reflect the unique requirements of the public sector. We analyse more than 1,300 benchmark papers for these criteria using an expert-validated LLM-assisted pipeline. Our results show that no single benchmark meets all of the criteria. Our findings provide a call to action for both researchers to develop public sector-relevant benchmarks and for public-sector officials to apply these criteria when evaluating their own agentic use cases.

LLM代理评测基准公共部门AI治理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。