arXiv:2505.23667cs.AI2025-05被引 3

用公式驱动强化学习,让大模型更准地做复杂表格的数值推理。

Formula-R1: Incentivizing LLM Reasoning over Complex Tables with Numerical Computation via Formula-Driven Reinforcement Learning

  • 通过生成可执行公式并用奖励信号训练模型,减少对标注数据依赖。
  • 在7个基准上显著提升复杂表格多步计算的准确率。
  • 适合需要精准数值推理的智能系统研究者和开发者。

表格是组织与分析数据的基础媒介,表推理对智能系统至关重要。尽管大语言模型具备强大通用推理能力,但在复杂表格上的数值推理仍表现不佳,尤其超出简单关系查询的场景。电子表格公式提供了强大且表达力强的可执行符号操作接口,能支持丰富推理模式,但现有大模型对此探索不足。本文提出Formula-R1,基于公式调优(Fortune)——一种公式驱动的强化学习框架,用于表推理。该方法通过执行成功率与答案正确性作为奖励信号,训练大模型生成可执行的电子表格公式以回答表中问题,从而降低对人工公式标注的依赖。我们在七个表推理基准上进行了广泛实验,结果表明,Formula Tuning显著提升了大模型在复杂表格及多步数值计算任务中的性能。此外,Formula-R1在受控对比实验中始终优于已有方法。除实证收益外,我们的深入分析揭示了强化学习在公式驱动表推理中的作用,凸显公式驱动强化学习提升大模型推理能力的广阔潜力。

原文摘要 · Abstract (English)

Tables are a fundamental medium for organizing and analyzing data, making table reasoning a critical capability for intelligent systems. Although large language models (LLMs) exhibit strong general reasoning abilities, they still struggle with accurate numerical reasoning over tabular data, particularly in complex table settings beyond simple relational lookup. Spreadsheet formulas provide a powerful and expressive interface for executable symbolic operations, enabling rich reasoning patterns that remain largely underexplored by existing LLMs. In this paper, we introduce Formula-R1, a model trained via Formula Tuning (Fortune), a formula-driven reinforcement learning (RL) framework for table reasoning. Formula Tuning trains LLMs to generate executable spreadsheet formulas for question answering over general tabular data, using execution success and answer correctness as reward signals, thereby reducing reliance on supervised formula annotations. We demonstrate the effectiveness of Formula Tuning through extensive experiments on seven table reasoning benchmarks. It substantially improves LLM performance on table reasoning, particularly for tasks involving complex tables and multi-step numerical computation. Moreover, Formula-R1 consistently outperforms prior methods under controlled comparison settings. Beyond empirical gains, our extensive analyses provide insights into the role of RL in formula-driven table reasoning, highlighting the broader potential of formula-driven RL to enhance reasoning capabilities in LLMs.

表格推理强化学习数值计算大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。