arXiv:2606.20700cs.MAcs.AI2026-06被引 1

让政策控制器可解释、可挑战、可迭代,提升监管模拟的透明度与适应性。

Machine-Coached Policy Revision in Adaptive Agent-Based Regulatory Simulation: A Controller-Level Contestability Layer

论文配图:Machine-Coached Policy Revision in Adaptive Agent-Based Regulatory Simulation: A Controller-Level Contestability Layer
图 1 · 摘自论文原文
  • 用可撤销规则表示政策决策,支持冲突识别与优先级管理。
  • 在排放调控实验中,通过新增宽松规则降低过度保守问题复发率。
  • 适合关注政策可解释性与动态调优的研究者或监管设计者。

以政策为导向的基于代理的模型正被广泛用于复杂自适应社会技术系统的监管干预研究。现有自适应ABM框架区分了静态与自适应代理、固定与可变政策及不同控制器设计,但多数诊断流程仍为事后分析:仿真轨迹在结束后才被分析,结果未系统反馈至政策控制器。本文提出一种轻量级的机器辅助政策修订层,将政策决策表示为带有显式冲突与优先级的可撤销规则,生成控制器行为解释,并将诊断失败转化为规则增删或优先级调整。该贡献并非新最优控制器,也不承诺对任意机器辅导提供形式化保证。其核心在于实现控制器层面的可争辩性:政策决策可在保留仿真条件下被解释、质疑、修订与重新评估。实验采用简化的排放调控ABM,控制实验聚焦于VPVA制度下的过度保守问题。预设辅导模板向符号控制器添加松弛规则,在保留违规、超调与波动阈值的前提下,显著降低了持留种子下过度保守问题的复发率。论文主张,机器辅导应视为可解释自适应ABM的控制器层面延伸,与因果、信息论及轨迹诊断方法形成互补。

原文摘要 · Abstract (English)

Policy-oriented agent-based models are increasingly used to study regulatory interventions in complex adaptive socio-technical systems. Recent adaptive ABM frameworks distinguish between static and adaptive agents, fixed and adaptive policies, and alternative controller designs. However, most diagnostic workflows remain ex post: trajectories are analysed after simulation, but the resulting evidence is not systematically fed back into the policy controller. This paper proposes a lightweight machine-coached policy-revision layer for adaptive agent-based regulation. The layer represents policy decisions as defeasible rules with explicit conflicts and priorities, generates explanations for controller actions, and allows diagnostic failures to be translated into rule additions, removals, or priority changes. The contribution is not a new optimal controller and does not claim formal guarantees for unrestricted machine coaching. Instead, it provides a simulation-compatible operationalization of controller-level contestability: policy decisions can be explained, challenged, revised, and re-evaluated in held-out simulation runs. A stylized emissions-regulation ABM is used as the experimental component. A controlled simulation experiment focuses on an over-conservatism failure in the VPVA regime. The predefined coaching template adds a relaxation rule to the symbolic controller, reducing over-conservatism recurrence under held-out seeds while preserving violation, overshoot, and volatility guardrails. The paper argues that machine coaching is best understood as a controller-level extension of explainable adaptive ABM, complementary to causal, information-theoretic, and trajectory-based diagnostics.

政策模拟可解释性自适应系统机器辅导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。