arXiv:2507.16864cs.LGcs.AI2025-07

用双曲空间增强强化学习,让模型更高效解决多步推理问题。

Reinforcement Learning in hyperbolic space for multi-step reasoning

  • 将双曲嵌入融入Transformer的强化学习框架,捕捉层次化结构
  • 在数学和控制任务上准确率提升32%~45%,计算耗时减少16%~32%
  • 适合处理具有层级关系的复杂推理任务,如数学证明与决策规划

多步推理是人工智能的核心挑战,涵盖数学求解到动态环境决策。强化学习(RL)通过优化长期奖励展现潜力,但传统方法在信用分配、高维状态表示和稳定性方面存在瓶颈。本文提出一种新框架,将双曲Transformer融入强化学习,利用双曲嵌入有效建模层次结构。通过理论分析、算法设计与实验验证,结果表明:在Frontier Math基准上,准确率提升32%~44%,计算时间减少16%~32%;在非线性最优控制任务上,准确率提升43%~45%,计算时间减少16%~17%。研究展示了双曲Transformer在多步推理任务中的巨大潜力。

原文摘要 · Abstract (English)

Multi-step reasoning is a fundamental challenge in artificial intelligence, with applications ranging from mathematical problem-solving to decision-making in dynamic environments. Reinforcement Learning (RL) has shown promise in enabling agents to perform multi-step reasoning by optimizing long-term rewards. However, conventional RL methods struggle with complex reasoning tasks due to issues such as credit assignment, high-dimensional state representations, and stability concerns. Recent advancements in Transformer architectures and hyperbolic geometry have provided novel solutions to these challenges. This paper introduces a new framework that integrates hyperbolic Transformers into RL for multi-step reasoning. The proposed approach leverages hyperbolic embeddings to model hierarchical structures effectively. We present theoretical insights, algorithmic details, and experimental results that include Frontier Math and nonlinear optimal control problems. Compared to RL with vanilla transformer, the hyperbolic RL largely improves accuracy by (32%~44%) on FrontierMath benchmark, (43%~45%) on nonlinear optimal control benchmark, while achieving impressive reduction in computational time by (16%~32%) on FrontierMath benchmark, (16%~17%) on nonlinear optimal control benchmark. Our work demonstrates the potential of hyperbolic Transformers in reinforcement learning, particularly for multi-step reasoning tasks that involve hierarchical structures.

强化学习多步推理双曲几何Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。