arXiv:2410.21741cs.CLcs.AI2024-10中稿 · ICAIF 24被引 25

用多智能体反思机制提升金融问答中的数值推理能力

Enhancing Financial Question Answering with a Multi-Agent Reflection Framework

  • 引入批判性智能体对推理过程和答案进行分步反思
  • 使开源模型在金融问答上性能提升15%(LLaMA3-8B)
  • 适合需要低成本高精度金融分析的开发者与研究者

尽管大语言模型在自然语言处理任务中表现优异,但在涉及数值推理的金融问答任务中仍存在困难。近期基于多智能体的框架在多步推理中展现出显著效果,这正是金融问答所需的关键能力——从表格和文本中提取信息,并对数据进行数值推理以得出答案。本文提出一种包含批判性智能体的多智能体框架,该智能体会反思每个问题的推理步骤和最终答案。此外,我们还引入多个专注于不同答案维度的批判性智能体以进一步增强系统。实验结果表明,该框架相比单智能体推理显著提升了性能,其中LLaMA3-8B模型平均性能提升15%,而LLaMA3-70B模型提升5%。此外,该框架在多数情况下达到甚至超过更大的单智能体模型如LLaMA3.1-405B和GPT-4o-mini的表现,仅略逊于Claude-3.5 Sonnet。总体而言,该框架为提升开源大模型在金融问答中的表现提供了一种高效且经济的解决方案。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) have shown impressive capabilities in numerous Natural Language Processing (NLP) tasks, they still struggle with financial question answering (QA), particularly when numerical reasoning is required. Recently, LLM-based multi-agent frameworks have demonstrated remarkable effectiveness in multi-step reasoning, which is crucial for financial QA tasks as it involves extracting relevant information from tables and text and then performing numerical reasoning on the extracted data to infer answers. In this study, we propose a multi-agent framework incorporating a critic agent that reflects on the reasoning steps and final answers for each question. Additionally, we enhance our system by adding multiple critic agents, each focusing on a specific aspect of the answer. Our results indicate that this framework significantly improves performance compared to single-agent reasoning, with an average performance increase of 15% for the LLaMA3-8B model and 5% for the LLaMA3-70B model. Furthermore, our framework performs on par with, and in some cases surpasses, larger single-agent LLMs such as LLaMA3.1-405B and GPT-4o-mini, though it falls slightly short compared to Claude-3.5 Sonnet. Overall, our framework presents an effective solution to enhance open-source LLMs for financial QA tasks, offering a cost-effective alternative to larger models like Claude-3.5 Sonnet.

金融问答多智能体数值推理反思机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。