为复杂卡牌游戏设计因果强化学习基准,支持因果归因与策略可审计性。
Causal Reinforcement Learning for Complex Card Games: A Magic The Gathering Benchmark

- 基于魔法规则构建带显式因果结构的强化学习环境
- 新算法在部分观测下实现可比胜率并提升因果归因精度
- 适合研究因果推理、世界模型与大模型代理的学者使用
因果强化学习缺乏结合序列决策、隐藏信息、大规模掩码动作空间及明确因果结构的复杂系统基准。我们提出MTG-Causal-RL,一个基于《魔法风云会》构建的Gymnasium基准,包含3077维部分观测、478维掩码离散动作空间、五种标准套牌、三种奖励机制,以及手写结构因果模型(SCM)对战略变量的关系建模。每轮实验暴露因果变量、SCM预测干预效果及各因子贡献轨迹,使因果归因、留一法跨套牌迁移和策略可审计性成为核心评估指标。我们引入基准方法:随机策略、启发式策略、掩码PPO、因果世界模型变体PPO及匹配架构的标量控制。提出因果图因子化优势PPO(CGFA-PPO),利用获胜概率的SCM父节点作为对齐因子的批评者目标,并引入干预校准损失。所有对比均采用配对种子、配对自助置信区间及霍尔姆-博尼费罗尼校正。掩码PPO与CGFA-PPO在分布内达到竞争性胜率,显著优于随机基线;因子校准轨迹与留一迁移差距揭示了标量胜率无法捕捉的诊断结构。我们公开发布基准、基准结果与完整评估协议。通过将策略丰富、部分可观测的领域与显式因果接口及统计协议结合,MTG-Causal-RL为因果强化学习、世界模型与大模型代理研究提供了共同测试平台,可回答现有基准无法联合提出的三类问题:掩码动作空间下的因果归因、套牌间的结构迁移,以及基于SCM的策略审计。
原文摘要 · Abstract (English)
Causal reinforcement learning (RL) lacks benchmarks for complex systems that combine sequential decision making, hidden information, large masked action spaces, and explicit causal structure. We introduce MTG-Causal-RL, a Gymnasium benchmark built on Magic: The Gathering with a 3,077-dimensional partial observation, a 478-action masked discrete action space, five competitive Standard archetypes, three reward schemes, and a hand-specified Structural Causal Model (SCM) over strategic variables. Every episode exposes causal variables, SCM-predicted intervention effects, and per-factor credit traces, making causal credit assignment, leave-one-out cross-archetype transfer, and policy auditability first-class metrics. We adapt a panel of reference baselines: random, heuristic, masked PPO, a causal-world-model PPO variant, and an architecture-matched scalar control. We propose Causal Graph-Factored Advantage PPO (CGFA-PPO) as a reference causal agent that uses SCM parents of win probability as factor-aligned critic targets with an intervention-calibration loss. All comparisons use paired seeds, paired-bootstrap confidence intervals, and Holm-Bonferroni correction within pre-registered families. Masked PPO and CGFA-PPO reach competitive in-distribution win rates and exceed the random baseline; per-factor calibration trajectories and leave-one-out transfer gaps expose diagnostic structure that scalar win rate alone cannot. We release the benchmark, reference-baseline results, and full evaluation protocol openly. By coupling a strategically rich, partially observed domain with an explicit causal interface and statistical protocol, MTG-Causal-RL gives causal-RL, world-model, and LLM-agent research a shared testbed for questions current benchmarks cannot pose together: causal credit assignment under masked action spaces, structural transfer across archetypes, and SCM-grounded policy auditability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。