用大模型构建金融问答流水线,提升决策信息提取效率。
FinQAPT: Empowering Financial Decisions with End-to-End LLM-driven Question Answering Pipeline
- 构建端到端金融问答流水线,整合文档检索与大模型推理
- 在FinQA数据集上实现80.6%的模块级准确率,但整体性能下降
- 提出动态提示与聚类负采样,助力数值问答与上下文抽取
金融决策依赖于对海量金融文档中相关信息的分析。为应对这一挑战,我们开发了FinQAPT——一个端到端的问答流水线,可基于查询自动识别相关财务报告、提取关键上下文,并利用大语言模型(LLMs)完成下游任务。为评估该流水线,我们在FinQA数据集上测试了多种优化方法,引入了一种新型基于聚类的负采样技术以增强上下文提取,并提出一种名为动态N-shot提示的新方法以提升LLMs在数值问答上的表现。模块层面,达到80.6%的准确率,处于领先水平;但在流水线整体层面,因从财务报告中提取相关上下文存在困难,性能有所下降。我们对各模块及端到端流程进行了详细错误分析,指出了需解决的具体挑战,以实现对复杂金融任务的稳健处理。
原文摘要 · Abstract (English)
Financial decision-making hinges on the analysis of relevant information embedded in the enormous volume of documents in the financial domain. To address this challenge, we developed FinQAPT, an end-to-end pipeline that streamlines the identification of relevant financial reports based on a query, extracts pertinent context, and leverages Large Language Models (LLMs) to perform downstream tasks. To evaluate the pipeline, we experimented with various techniques to optimize the performance of each module using the FinQA dataset. We introduced a novel clustering-based negative sampling technique to enhance context extraction and a novel prompting method called Dynamic N-shot Prompting to boost the numerical question-answering capabilities of LLMs. At the module level, we achieved state-of-the-art accuracy on FinQA, attaining an accuracy of 80.6%. However, at the pipeline level, we observed decreased performance due to challenges in extracting relevant context from financial reports. We conducted a detailed error analysis of each module and the end-to-end pipeline, pinpointing specific challenges that must be addressed to develop a robust solution for handling complex financial tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。