首次系统研究生成式过程奖励模型在表格问答中的应用
Exploring Generative Process Reward Modeling for Semi-Structured Data: A Case Study of Table Question Answering
- 用文本与代码双重验证提升推理步骤评分
- 结合步骤评估与答案正确率,发现二者相关性弱
- 适合关注大模型推理可解释性与验证机制的研究者
过程奖励模型(PRMs)通过逐步评估候选解并基于累积步骤得分选择答案,提升了大语言模型在复杂推理任务中的表现。尽管在数学等领域效果显著,其在半结构化数据任务如表格问答(TQA)中的适用性尚未被探索。TQA面临信息冗余、推理步骤松散连接及领域特异性推理等挑战。本文首次对TQA任务中前沿生成式PRMs进行了系统研究,从答案和步骤两个角度进行评估。结果表明,结合文本与代码验证的PRMs虽能辅助解的选择,但难以泛化到域外数据。分析显示,步骤级验证性能与答案准确率之间存在较弱相关性,可能源于推理步骤间依赖性弱、因果关联松散。研究揭示了当前PRMs在TQA中的局限性,并为构建更鲁棒的流程感知验证器提供了重要启示。
原文摘要 · Abstract (English)
Process reward models (PRMs) enhance complex reasoning in large language models (LLMs) by evaluating candidate solutions step-by-step and selecting answers based on aggregated step scores. While effective in domains such as mathematics, their applicability to tasks involving semi-structured data, like table question answering (TQA), remains unexplored. TQA poses unique challenges for PRMs, including abundant irrelevant information, loosely connected reasoning steps, and domain-specific reasoning. This work presents the first systematic study of PRMs for TQA. We evaluate state-of-the-art generative PRMs on TQA from both answer and step perspectives. Results show that PRMs that combine textual and code verification can aid solution selection but struggle to generalize to out-of-domain data. Analysis reveals a weak correlation between performance in step-level verification and answer accuracy, possibly stemming from weak step dependencies and loose causal links. Our findings highlight limitations of current PRMs on TQA and offer valuable insights for building more robust, process-aware verifiers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。