研究政策变化下多智能体系统的迁移学习,发现结构知识可提升小样本表现。
Transfer Learning Across Policy Regimes in Adaptive Multi-Agent Systems

- 将政策变迁建模为迁移学习问题,利用历史结构知识约束新环境假设空间。
- 当政策与结果呈线性单调关系时,迁移学习显著提升小样本性能;引入阈值突变则导致负迁移。
- 适用于需应对政策调整的智能系统设计者,强调稳定结构可复用,关系改变则需谨慎。
政策模型常假设政策工具与结果之间的关系在不同制度条件下保持稳定。但在自适应社会技术系统中,这种假设可能失效:监管变化会改变激励机制,代理人可能策略性响应,政策变量到总体结果的映射可能发生改变。本文将此类制度变迁视为自适应多智能体系统中的迁移学习问题。一个政策制度被表示为由可观测输入分布和目标函数(将政策变量映射到结果)所诱导的学习任务。我们比较了两种学习器:一种是空白初值学习器,在新制度下搜索灵活假设空间;另一种是迁移学习器,其有效假设空间受前一制度的结构知识约束。当该约束能保留新目标函数但降低有效复杂度时,迁移有益;若约束排除了新目标函数,则造成误设,导致迁移有害。基于简化的排放监管实验环境与动态代理模型鲁棒性实验,结果表明:当目标制度保持线性单调的税收-排放关系时,迁移可提升实证小样本性能;而当目标制度引入阈值断裂时,相同转移结构引发负迁移——外推误差持续高位,在线预测错误增多,重复在线流中累积误差与最终窗口误差均更大。方法论贡献在于:先前监管经验应被重用于捕捉稳定的结构不变量,但当政策-结果关系发生改变时需谨慎处理。
原文摘要 · Abstract (English)
Policy models often assume that the relationship between a policy instrument and its outcome remains stable across institutional conditions. In adaptive socio-technical systems this assumption may fail: regulatory change can alter incentives, agents can respond strategically, and the mapping from policy variables to aggregate outcomes can change. This paper studies such regime change as a transfer-learning problem in adaptive multi-agent systems. A policy regime is represented as a learning problem induced by an observable input distribution and a target function mapping policy variables to outcomes. We compare a blank-slate learner that searches a flexible hypothesis class in the new regime with a transfer learner whose effective hypothesis class is restricted by structural knowledge from the previous regime. Transfer is beneficial when this restriction preserves the new target function while reducing effective complexity; it is harmful when the restriction excludes the new target and creates misspecification. A stylized emissions-regulation experimental environment and a dynamic ABM robustness experiment support the claim. When the target regime preserves an affine monotone tax-emissions relation, transfer improves empirical small-sample performance. When the target regime introduces a threshold break, the same transferred structure produces negative transfer: held-out error remains high, online prediction generates more mistakes, and repeated online streams show larger cumulative and final-window error under misspecification. The contribution is methodological: previous regulatory experience should be reused when it captures stable structural invariants, but treated cautiously when policy change alters the policy-outcome relationship.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。