arXiv:2602.11829cs.LGcs.GT2026-02中稿 · ICLR

用对手塑造算法改善投资决策,让市场更可持续

Towards Sustainable Investment Policies Informed by Opponent Shaping

  • 引入优势对齐算法调整投资者学习过程
  • 在模拟中实现比原模型更高合作率的长期收益
  • 为气候政策提供可落地的激励机制设计参考

应对气候变化需要全球协调,但理性经济主体常优先短期利益,导致社会困境。InvestESG 是一个捕捉投资者与企业间动态互动的多智能体仿真模型。本文形式化刻画了该模型中跨期社会困境的条件,推导出个体激励偏离集体福利的理论阈值。在此基础上,应用一种可扩展的对手塑造算法——优势对齐(Advantage Alignment),影响 InvestESG 中智能体的学习过程。我们提供了理论解释:优势对齐通过偏移学习动态,系统性倾向合作均衡。实验结果表明,战略性地塑造经济主体的学习过程,能显著改善长期结果,为政策设计提供依据,使市场激励更契合长期可持续目标。

原文摘要 · Abstract (English)

Addressing climate change requires global coordination, yet rational economic actors often prioritize immediate gains over collective welfare, resulting in social dilemmas. InvestESG is a recently proposed multi-agent simulation that captures the dynamic interplay between investors and companies under climate risk. We provide a formal characterization of the conditions under which InvestESG exhibits an intertemporal social dilemma, deriving theoretical thresholds at which individual incentives diverge from collective welfare. Building on this, we apply Advantage Alignment, a scalable opponent shaping algorithm shown to be effective in general-sum games, to influence agent learning in InvestESG. We offer theoretical insights into why Advantage Alignment systematically favors socially beneficial equilibria by biasing learning dynamics toward cooperative outcomes. Our results demonstrate that strategically shaping the learning processes of economic agents can result in better outcomes that could inform policy mechanisms to better align market incentives with long-term sustainability goals.

多智能体可持续投资策略塑造

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。