在量化推理限制下,用多帧和多策略头突破博弈游戏的表征瓶颈。
Mastering NIM and Impartial Games with Weak Neural Networks: An AlphaZero-inspired Multi-Frame Approach
- 设计多策略头与多帧架构,突破单帧推理的线性局限
- 两帧模型实现接近完美的状态恢复准确率,多头模型完全分类胜负位置
- 适合研究低精度强化学习中结构先验的重要性
我们在固定延迟、固定尺度量化推理(FSQI)环境下研究公平博弈。在此固定尺度、有界范围的设定中,我们证明推理可由常数深度、多项式规模的布尔电路(AC0)模拟,由此产生最坏情况下的表征障碍:单帧智能体无法强攻NIM游戏,因为最优策略依赖全局尼姆和(奇偶性)。在确定性展开接口下,单一展开策略头仅能揭示不变量的一个固定线性函数,单纯增加展开预算无法弥补缺失信息。我们提出两种结构化解法:(1) 多策略头展开架构,通过不同展开通道恢复完整不变量;(2) 多帧架构,追踪局部尼姆数差异并支持状态恢复。多个设置的实验结果与此预测一致:单头基线表现接近随机,两帧模型达到近完美的恢复准确率,多头FSM控制对局实现完美胜负位置分类。总体结果支持在FSQI/AC0框架下,显式结构先验(历史/差值或多重展开通道)至关重要。
原文摘要 · Abstract (English)
We study impartial games under fixed-latency, fixed-scale quantised inference (FSQI). In this fixed-scale, bounded-range regime, we prove that inference is simulable by constant-depth polynomial-size Boolean circuits (AC0). This yields a worst-case representational barrier: single-frame agents in the FSQI/AC0 regime cannot strongly master NIM, because optimal play depends on the global nim-sum (parity). Under our stylised deterministic rollout interface, a single rollout policy head from the structured family analysed here reveals only one fixed linear functional of the invariant, so increasing rollout budget alone does not recover the missing bits. We derive two structural bypasses: (1) a multi-policy-head rollout architecture that recovers the full invariant via distinct rollout channels, and (2) a multi-frame architecture that tracks local nimber differences and supports restoration. Experiments across multiple settings are consistent with these predictions: single-head baselines stay near chance, while two-frame models reach near-perfect restoration accuracy and multi-head FSM-controlled shootouts achieve perfect win/loss position classification. Overall, the empirical results support the view that explicit structural priors (history/differences or multiple rollout channels) are important in the FSQI/AC0 regime.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。