arXiv:2411.18915cs.LGcs.CL2024-11被引 1

用弱监督训练小模型,让文档表格推理更高效

MATATA: Weakly Supervised End-to-End MAthematical Tool-Augmented Reasoning for Tabular Applications

  • 用最终结果做弱监督,无需逐步标注中间过程
  • 在FinQA和TAT-QA上达到开源小模型最佳性能
  • 适合想低成本部署智能推理系统的开发者

商业文档常包含大量含数值的表格与文本信息,需数学推理才能有效理解。尽管小型语言模型(SLMs)在此任务上仍表现不佳,工具增强的多步智能体虽表现更好,但依赖闭源大模型、外部数据或大量提示工程。本文提出MATATA,一种新型弱监督端到端方法,用于训练适用于文档表格应用的多步推理语言智能体。MATATA采用无标注范式,增强3.8B/8B规模的SLMs。其两阶段训练中,以多步推理链的最终结果作为弱监督信号,避免对每一步中间推理进行单独标注。通过自适应规划器和跨数据集共享工具,MATATA展现鲁棒性能。实验表明,MATATA在FinQA上达到当前最优,在基于开源SLM的推理方法中于TAT-QA上领先。尽管基于SLMs,其在TabMWP上的表现接近GPT-4框架。该弱监督方法实现无需中间标注的端到端多步推理智能体训练,为低成本、强能力的智能体系统发展提供支持。

原文摘要 · Abstract (English)

Business documents often contain substantial tabular and textual information with numerical values, requiring mathematical reasoning for effective document understanding. While Small Language Models (SLMs) still struggle at this task, tool-augmented multi-step agents perform better, at the cost of relying on closed-source or larger models, external data, or extensive prompt-engineering. This work introduces MATATA, a novel weakly supervised end-to-end approach to train multi-step reasoning language agents for document tabular applications. MATATA presents an annotation-free paradigm for each agent to enhance 3.8B/8B SLMs. During its two-stage training, MATATA uses the final outcome of the multi-step reasoning chain as weak supervision. This approach avoids having to individually supervise each intermediate agent in the reasoning chain. By employing an adaptive planner and shared tools across different datasets, MATATA shows robust performance. Experiments demonstrate that MATATA achieves state-of-the-art on FinQA, and on TAT-QA among reasoning methods based on open-source SLMs. Although being SLM-based, MATATA closely matches GPT-4-based frameworks on TabMWP. This novel weakly supervised approach enables training an end-to-end multi-step reasoning agent without intermediate supervision, supporting future developments of cost-effective powerful agentic systems.

数学推理小模型弱监督智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。