arXiv:2601.04695cs.AIcs.LG2026-01

提出Tape基准,测试强化学习在规则变化下的泛化能力。

Tape: A Cellular Automata Benchmark for Evaluating Rule-Shift Generalization in Reinforcement Learning

  • 设计隔离动态规则变化的可控环境,保持观测与动作接口不变。
  • 发现所有基线模型在分布外任务中性能显著下降,且对规则类型敏感。
  • 提供可比参考基准,适合评估鲁棒性与机制推理能力。

强化学习中的分布外泛化难以诊断,因基准变化常混杂动力学、观测、目标与奖励等多重因素。本文提出Tape,一个受控基准,仅分离潜在规则变化的动力学成分,同时固定观测-动作接口。该协议包含确定性分组、20次种子重复、自助置信区间报告,以及稀疏成功场景下的连续评估指标。在多个基线方法中,均观察到从内部分布(ID)到分布外(OOD)的性能持续下降,且在稳定、周期与混沌规则间存在显著异质性。值得注意的是,这种脆弱性甚至出现在刻意简化的1维确定性设定中,表明当前多数RL算法在最小干扰下仍对潜在规律变化极为敏感。为校准严格成功标准,报告了匹配协议的真实动态随机射击基准(p_oracle ≈ 0.187),并定义归一化得分ON(p) = 100p / p_oracle,作为预算受限的操作参考,而非全局最优界限。较小可行性规模(L = H = 16)下规则全可解,有助于区分可达性极限与策略失败。这些结果使Tape成为机制导向的鲁棒适应与潜在机制推断诊断工具,并为更广泛的通用人工智能评估提供受控基准,不作强通用智能充分性声明。

原文摘要 · Abstract (English)

Out-of-distribution generalization in reinforcement learning is hard to diagnose when benchmark shifts mix dynamics, observations, goals, and rewards. We address this with Tape, a controlled benchmark that isolates latent rule-shift in dynamics while keeping the observation-action interface fixed. The protocol combines deterministic splits, 20-seed replication, bootstrap uncertainty reporting, and continuous metrics for sparse-success regimes. Across baseline families, we find a consistent ID-to-OOD drop and strong heterogeneity across stable/periodic/chaotic rules. Importantly, this fragility appears even in an intentionally simple 1D deterministic setting, suggesting that many current RL algorithms remain brittle to latent-law changes under minimal confounds. To calibrate strict success, we report a protocol-matched true-dynamics random-shooting reference (p_oracle is almost 0.187) and oracle-normalized scores ON(p) = 100 p / p_oracle; this is a budgeted operational reference, not a global-optimality bound. A smaller feasibility regime (L = H = 16) with 100% rule-wise solvability helps separate reachability limits from policy failure. These results position Tape as a mechanism-oriented diagnostic for robust adaptation and latent-mechanism inference, and as a controlled benchmark relevant to broader AGI-oriented evaluation without making strong AGI sufficiency claims.

强化学习泛化能力基准测试规则迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。