arXiv:2506.08726cs.CLcs.AI2025-06被引 6

改进的智能体框架提升金融文档问答中的数值计算能力

Improved LLM Agents for Financial Document Question Answering

  • 引入计算器智能体与优化批评者智能体协同工作
  • 在无真实答案条件下性能优于现有最先进方法
  • 适合需要高精度财务分析的金融从业者使用

大型语言模型(LLMs)在众多自然语言处理任务中表现出色,但在包含表格和文本数据的金融文档数值问答任务中仍表现不佳。已有研究证明,在提供真实标签(oracle labels)的情况下,批评者智能体(即自我修正机制)有效。本文在此基础上,考察当缺乏真实标签时传统批评者智能体的表现,实验显示其性能显著下降。为此,本文提出一种改进的批评者智能体,并引入计算器智能体,该组合方法超越了先前最先进的「思维链」(program-of-thought)方法,且更安全可靠。此外,本文还研究了两类智能体之间的交互机制及其对整体性能的影响。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown impressive capabilities on numerous natural language processing tasks. However, LLMs still struggle with numerical question answering for financial documents that include tabular and textual data. Recent works have showed the effectiveness of critic agents (i.e., self-correction) for this task given oracle labels. Building upon this framework, this paper examines the effectiveness of the traditional critic agent when oracle labels are not available, and show, through experiments, that this critic agent's performance deteriorates in this scenario. With this in mind, we present an improved critic agent, along with the calculator agent which outperforms the previous state-of-the-art approach (program-of-thought) and is safer. Furthermore, we investigate how our agents interact with each other, and how this interaction affects their performance.

金融问答大模型智能体系统数值推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。