arXiv:2604.12995cs.CLcs.CY2026-04ACL

首个跨中美政策理解基准,评估大模型政策推理能力

PolicyLLM: Towards Excellent Comprehension of Public Policy for Large Language Models

  • 构建跨系统政策基准,按认知层次分三类任务
  • 专家模型在政策应用任务上准确率最高,结构化推理表现突出
  • 揭示现有大模型在政策理解上的短板,适合政策科技研究者

大型语言模型(LLMs)正日益融入现实决策,包括公共政策领域。然而,其对政策相关内容的理解与推理能力仍待深入探索。为填补这一空白,我们提出首个大规模跨系统(中美)政策理解评估基准 extbf{ extit{PolicyBench}},涵盖21,000个案例,覆盖广泛政策领域,反映真实治理的多样性和复杂性。基于布卢姆分类法,该基准评估三项核心能力:(1) extbf{记忆}:政策知识的事实性回忆;(2) extbf{理解}:概念与语境推理;(3) extbf{应用}:在真实政策场景中的问题解决。在此基础上,我们进一步提出 extbf{ extit{PolicyMoE}},一种面向政策领域的专家混合(MoE)模型,其专家模块分别对应上述认知层级。实验表明,所提模型在应用导向任务上表现优于记忆或概念理解任务,并在结构化推理任务中取得最高准确率。结果揭示了当前大模型在政策理解方面的关键局限,为构建更可靠、聚焦政策的模型指明方向。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly integrated into real-world decision-making, including in the domain of public policy. Yet, their ability to comprehend and reason about policy-related content remains underexplored. To fill this gap, we present \textbf{\textit{PolicyBench}}, the first large-scale cross-system benchmark (US-China) evaluating policy comprehension, comprising 21K cases across a broad spectrum of policy areas, capturing the diversity and complexity of real-world governance. Following Bloom's taxonomy, the benchmark assesses three core capabilities: (1) \textbf{Memorization}: factual recall of policy knowledge, (2) \textbf{Understanding}: conceptual and contextual reasoning, and (3) \textbf{Application}: problem-solving in real-life policy scenarios. Building on this benchmark, we further propose \textbf{\textit{PolicyMoE}}, a domain-specialized Mixture-of-Experts (MoE) model with expert modules aligned to each cognitive level. The proposed models demonstrate stronger performance on application-oriented policy tasks than on memorization or conceptual understanding, and yields the highest accuracy on structured reasoning tasks. Our results reveal key limitations of current LLMs in policy understanding and suggest paths toward more reliable, policy-focused models.

政策理解大模型基准测试MoE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。