用强化学习提升Transformer在多项式分解中的效率与精度
Discovering Hidden Algebraic Structures via Transformers with Rank-Aware Beam GRPO
- 提出面向代数问题的秩感知强化学习方法BGRPO
- 模型精度提升,推理计算量降低75%
- 适合需要高效符号计算的研究者与工程师
近期研究拓展了Transformer在逻辑推理与符号计算方面的能力。本文聚焦于函数分解中的非线性隐含模式发现,针对多变量多项式分解这一广泛应用于科学与工程的NP难问题。我们提出三方面贡献:第一,构建可精细控制复杂度的合成数据生成管道;第二,通过监督学习训练Transformer模型,并在缩放行为与泛化能力四个维度进行评估;第三,提出束搜索增强的相对策略优化(BGRPO),一种适用于复杂代数问题的秩感知强化学习方法。使用BGRPO微调后,模型准确率提升,束宽减少一半,推理计算量降低约75%。此外,模型在多项式简化任务中表现优异,部分情形下超越Mathematica。
原文摘要 · Abstract (English)
Recent efforts have extended the capabilities of transformers in logical reasoning and symbolic computations. In this work, we investigate their capacity for non-linear latent pattern discovery in the context of functional decomposition, focusing on the challenging algebraic task of multivariate polynomial decomposition. This problem, with widespread applications in science and engineering, is proved to be NP-hard, and demands both precision and insight. Our contributions are threefold: First, we develop a synthetic data generation pipeline providing fine-grained control over problem complexity. Second, we train transformer models via supervised learning and evaluate them across four key dimensions involving scaling behavior and generalizability. Third, we propose Beam Grouped Relative Policy Optimization (BGRPO), a rank-aware reinforcement learning method suitable for hard algebraic problems. Finetuning with BGRPO improves accuracy while reducing beam width by up to half, resulting in approximately 75% lower inference compute. Additionally, our model demonstrates competitive performance in polynomial simplification, outperforming Mathematica in various cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。