arXiv:2510.06217cs.AIcs.CL2025-10被引 12

让大模型更懂表格推理,用工具验证提升准确性

TaTToo: Tool-Grounded Thinking PRM for Test-Time Scaling in Tabular Reasoning

  • 基于表格操作设计可验证的思维链监督机制
  • 在5个基准上使推理模型准确率提升30.9%
  • 适合需要高精度表格分析的AI系统开发者

过程奖励模型(PRM)是增强大推理模型(LRM)推理能力的重要框架,尤其在测试时扩展(TTS)场景中表现突出。然而,现有PRM在表格推理任务中的应用仍不充分。我们通过实证分析发现,当前PRM在子表检索、模式交互等表格特有操作上表现不佳,成为性能瓶颈。为此,我们提出TaTToo——一种基于工具验证的表格导向PRM框架,具备显式处理表格推理步骤并融合工具验证以实现精准奖励监督的能力。我们构建了包含60,000+高质量步级标注的数据集,结合工具执行与表格验证逻辑。在此基础上,采用双阶段训练:先冷启动监督微调以捕捉工具使用模式,再通过工具引导的强化学习对齐表格验证目标。在5个涵盖数值推理、事实核验与数据分析的挑战性表格推理基准上,TaTToo使下游策略模型在推理阶段性能提升30.9%,仅用80亿参数即超越如Qwen-2.5-Math-PRM-72B等强大基线,且在多种TTS策略下表现出强泛化能力。

原文摘要 · Abstract (English)

Process Reward Models (PRMs) have recently emerged as a powerful framework for enhancing the reasoning capabilities of large reasoning models (LRMs), particularly in the context of test-time scaling (TTS). However, their potential for supervising LRMs on tabular reasoning domains remains underexplored. Through detailed empirical analyses, we identify that existing PRMs, though widely adopted for supervising text-only reasoning steps, struggle with table-specific operations such as sub-table retrieval and schema interaction, leading to critical performance bottlenecks. To address this limitation, we propose TaTToo, a novel table-grounded PRM framework that (i) reasons explicitly over tabular reasoning steps and (ii) integrates tool-based verification to provide precise reward supervision. Concretely, we first design a scalable data curation pipeline that constructs over 60k high-quality step-level annotations by integrating table verification rationales with tool-based executions. Building on the collected data, we train TaTToo with a dual-stage paradigm: cold-start supervised fine-tuning to capture tool-use reasoning patterns, followed by reinforcement learning with tool-grounded reward shaping to align our model with table-based verification. We provide a comprehensive evaluation of the policy improvement induced by our newly designed PRM. Across 5 challenging tabular reasoning benchmarks covering numerical reasoning, fact-checking, and data analysis, TaTToo improves downstream policy LRMs by 30.9% at inference, surpasses strong PRM baselines such as Qwen-2.5-Math-PRM-72B with only 8B parameters, and demonstrates strong generalizability across diverse TTS strategies.

表格推理强化学习工具使用大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。