提升大模型对数学表达式的精准推理能力,支持复杂格式输出。
Reasoning over mathematical objects: on-policy reward modeling and test time aggregation
- 构建了专门用于数学对象推导的训练数据集和评测基准Principia。
- 通过在线策略裁判训练,显著提升模型在数学表达式生成上的表现。
- 支持测试时计算资源扩展,实现推理结果聚合,适合科研与教育场景。
精确推导数学对象是下游科学应用(如数学、物理、化学)的核心需求,要求推理最终产出形式化表达式。然而当前大模型在数学与科学推理评估中仍依赖数值或选择题等简化答案格式,以方便自动评分。本文提出三项改进:(i) 构建并发布用于数学对象推导的训练数据与评测基准——Principia套件;(ii) 提供结合强语言模型裁判与验证器的训练方案,证明在线策略裁判训练能有效提升性能;(iii) 展示在线策略训练可用于测试时计算扩展,通过聚合实现性能提升。实验表明,Qwen3-235B与o3在Principia上表现不佳,而本文方法在多种大模型架构上均带来显著提升,并同步改善现有数值与多选题任务表现,证明推理能力具备跨格式泛化性。
原文摘要 · Abstract (English)
The ability to precisely derive mathematical objects is a core requirement for downstream STEM applications, including mathematics, physics, and chemistry, where reasoning must culminate in formally structured expressions. Yet, current LM evaluations of mathematical and scientific reasoning rely heavily on simplified answer formats such as numerical values or multiple choice options due to the convenience of automated assessment. In this paper we provide three contributions for improving reasoning over mathematical objects: (i) we build and release training data and benchmarks for deriving mathematical objects, the Principia suite; (ii) we provide training recipes with strong LLM-judges and verifiers, where we show that on-policy judge training boosts performance; (iii) we show how on-policy training can also be used to scale test-time compute via aggregation. We find that strong LMs such as Qwen3-235B and o3 struggle on Principia, while our training recipes can bring significant improvements over different LLM backbones, while simultaneously improving results on existing numerical and MCQA tasks, demonstrating cross-format generalization of reasoning abilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。