用双曲空间提升大模型推理,让搜索更高效准确
Latent Poincaré Shaping for Agentic Reinforcement Learning
- 在双曲潜空间中构建搜索树,利用负曲率扩展表达能力
- 在MATH-500上将模型准确率从66.0%提升至88.2%
- 支持测试时自引导扩展,适合需要强推理的数学求解任务
我们提出LaPha,一种在双曲潜空间中训练类似AlphaZero的大语言模型智能体的方法。在LaPha框架下,搜索过程可视为从提示根节点出发、向双曲球边界外生长的树结构,负曲率使得容量随半径指数级增长。通过使用双曲测地距离衡量规则验证正确性,定义节点势能,并基于势能差分配密集过程奖励。我们进一步在共享潜空间上附加轻量级价值头,实现几乎无额外开销的测试时自引导扩展。在MATH-500上,LaPha将Qwen2.5-Math-1.5B的准确率从66.0%提升至88.2%。采用价值头引导搜索后,LaPha-1.5B在AIME'24上达到56.7%准确率,LaPha-7B在AIME'24和AIME'25上分别达到60.0%和53.3%。
原文摘要 · Abstract (English)
We propose LaPha, a method for training AlphaZero-like LLM agents in a Poincaré latent space. Under LaPha, the search process can be visualized as a tree rooted at the prompt and growing outward from the origin toward the boundary of the Poincaré ball, where negative curvature provides exponentially increasing capacity with radius. Using hyperbolic geodesic distance to rule-verified correctness, we define a node potential and assign dense process rewards by potential differences. We further attach a lightweight value head on the same shared latent space, enabling self-guided test-time scaling with almost no additional overhead. On MATH-500, LaPha improves Qwen2.5-Math-1.5B from 66.0% to 88.2%. With value-head-guided search, LaPha-1.5B reaches 56.7% accuracy on AIME'24, and LaPha-7B further achieves 60.0% on AIME'24 and 53.3% on AIME'25.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。