arXiv:2509.06213cs.LGcs.AI2025-09

用隐藏规则游戏测试强化学习的推理能力

Toward a Metrology for Artificial Intelligence: Hidden-Rule Environments and Reinforcement Learning

  • 设计隐藏规则棋盘环境,让智能体推断规则并执行
  • 对象中心表征比特征中心更利于高效学习
  • 适合研究智能体推理与泛化能力的科研人员

我们研究了在「隐藏规则游戏」(GOHR)环境中的强化学习,这是一个复杂的谜题任务:智能体需推断并执行隐藏规则,通过将游戏棋子放入6×6棋盘上的桶中来清空棋盘。采用基于Transformer的A2C算法进行训练,探索特征中心(FC)和对象中心(OC)两种状态表示策略。智能体仅能获得部分观测,必须在经验中同时推断规则并学习最优策略。我们在多种基于规则和基于试验列表的实验设置下评估模型,分析迁移效果及表征对学习效率的影响。

原文摘要 · Abstract (English)

We investigate reinforcement learning in the Game Of Hidden Rules (GOHR) environment, a complex puzzle in which an agent must infer and execute hidden rules to clear a 6$\times$6 board by placing game pieces into buckets. We explore two state representation strategies, namely Feature-Centric (FC) and Object-Centric (OC), and employ a Transformer-based Advantage Actor-Critic (A2C) algorithm for training. The agent has access only to partial observations and must simultaneously infer the governing rule and learn the optimal policy through experience. We evaluate our models across multiple rule-based and trial-list-based experimental setups, analyzing transfer effects and the impact of representation on learning efficiency.

强化学习推理能力状态表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。