构建可验证的政策模拟基准,用真实数据评估大模型对政策影响的预测能力。
GPS-Bench: A Governance Policy Benchmark for Automating Policy Analysis

- 基于真实立法与披露数据重建政策相关方及其行为
- 发现微调训练比分解推理更有效,但分解能揭示作用机制
- 适合研究政策分析、多智能体系统与可信生成的学者
政策分析不仅需预测提案能否通过,还需识别受影响方、其反应及后续影响。现有基于大模型的政策模拟缺乏实证验证。我们提出GPS-Bench,一个基于真实证据的治理政策模拟基准,利用立法记录、游说披露、监管文件、企业财报、经济数据等公开信息,将政策与相关方、行动及下游影响关联起来。相关方从时间戳记录中重构,而非预设原型,确保角色有明确证据来源;人工标注形成黄金标准集,另一大模型基于检索证据标注的案例作为白银监督数据,不用于测试。所有推理模式共享同一接地状态与输出结构,使‘多智能体模拟是否有效’成为可控对比:我们比较联合推理、独立与通信型智能体、图方法及权重级微调在单一政策状态下的表现。在接地记录上微调获得最强的个体影响预测,分解虽未提升精度,但揭示了作用机制。智能体持有私有且非相同的证据,仅可见自身暴露条款,并向特定合作方提出具体联合提议,说明提供什么、需要什么以及协作优于单独行动的理由,从而可检验形成的联盟是否符合原始记录承诺。因此,GPS-Bench为研究证据、主体建模与多智能体交互如何提升政策结果预测与解释力提供了统一实证场景。
原文摘要 · Abstract (English)
Policy analysis requires more than predicting whether a proposal will pass: it requires identifying who will be affected, how those actors respond, and what follows. LLM-based policy simulations model these processes at scale, but their validity is hard to establish when plausible behaviour is never compared with observed outcomes. We introduce GPS-Bench, an evidence-grounded benchmark for governance policy simulation that links policies to relevant actors, actor actions and downstream impacts using legislative records, lobbying disclosures, regulatory documents, corporate filings, economic data and other public evidence. Actors are reconstructed from the dated record rather than prompted as archetypes, so a persona is an evidence object with provenance; a human-annotated pool forms the Gold evaluation set, while cases labelled by a separate LLM from retrieved evidence are treated as Silver supervision and never as test labels. Because every inference mode reads the same grounded state and emits the same schema, GPS-Bench turns "does multi-agent simulation help?" into a controlled comparison: we contrast joint reasoning, independent and communicating actor agents, graph-based methods and weight-level fine-tuning over one policy state. Fine-tuning on the grounded record gives the strongest actor-level impact prediction, and decomposition does not beat it; what decomposition adds is mechanism. Agents hold private, non-identical evidence, each seeing its own exposure clause, and address named partners with concrete joint proposals, what they offer, what they need in return, and why acting together beats acting alone, so the coalitions that form can be checked against the commitments the record holds. GPS-Bench therefore gives a common empirical setting for studying when evidence, actor modelling and multi-agent interaction improve the prediction and interpretation of policy outcomes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。