用分层强化学习找交换代数反例,突破稀疏奖励难题
Hierarchical Reinforcement Learning for Sparse-Reward Search in Commutative Algebra

- 基于约束选项的分层强化学习框架,结合等变图神经网络
- 在多种次数下均优于传统RL和贪心搜索,成功发现反例
- 首次将分层强化学习应用于交换代数领域,适合数学与AI交叉研究者
将机器学习应用于解决长期存在的数学猜想极具挑战性,主要因奖励极度稀疏。以卡拉伊代数哈希猜想为例,我们将构造其反例的问题重新建模为图上的稀疏奖励强化学习任务。提出一种带有约束选项的分层强化学习框架,采用等变图神经网络策略,能够有效学习该任务的时间抽象。在多种次数范围内进行评估,结果表明该方法持续优于经典强化学习算法及贪心搜索。通过利用问题的层次结构,首次实现了分层强化学习在交换代数领域的应用。
原文摘要 · Abstract (English)
Applying machine learning techniques to solving long-standing mathematical conjectures can be particularly challenging due to their extreme reward sparsity. As an illustrative example, we consider Kalai's algebraic Hirsch conjecture and recast the construction of its counterexamples as a sparse-reward reinforcement learning problem on graphs. We propose a constrained options-based HRL framework with an equivariant graph neural network policy, which allows us to learn useful temporal abstractions for this task. We evaluate our approach over a wide range of degrees and demonstrate that it consistently outperforms classical RL algorithms as well as greedy search. By exploiting the hierarchical structure of the problem, we effectively provide a first-of-its-kind application of HRL to a problem in commutative algebra.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。