构建金融分析基准数据集,评估大模型生成财报报告能力。
Towards Competent AI for Fundamental Analysis in Finance: A Benchmark Dataset and Evaluation
- 拆解财务分析为信息提取、指标计算、逻辑推理三步评估
- 发现大模型在指标计算上准确率超85%,但逻辑推理常出错
- 适合研究金融AI的开发者与量化投资机构参考
生成式AI,尤其是大型语言模型(LLMs),正开始通过自动化任务和解析复杂金融信息来改变金融行业。一个极具前景的应用是自动生成基本面分析报告,这对投资决策、信用风险评估、企业并购指导等至关重要。尽管当前LLMs尝试仅通过单个提示生成报告,但准确性风险显著。错误分析可能导致误投、监管问题和信任丧失。现有金融基准主要评估LLMs回答金融问题的能力,却未反映其在生成分析报告这类真实任务中的表现。本文提出FinAR-Bench,一个聚焦财务报表分析的核心能力基准数据集。为提升评估精度与可靠性,我们把该任务分解为三个可测量步骤:关键信息提取、财务指标计算和逻辑推理。这一结构化方法使我们能客观评估LLMs在各环节的表现。研究结果清晰揭示了当前大模型在基本面分析中的优势与局限,并提供了更贴近实际金融场景的评估路径。
原文摘要 · Abstract (English)
Generative AI, particularly large language models (LLMs), is beginning to transform the financial industry by automating tasks and helping to make sense of complex financial information. One especially promising use case is the automatic creation of fundamental analysis reports, which are essential for making informed investment decisions, evaluating credit risks, guiding corporate mergers, etc. While LLMs attempt to generate these reports from a single prompt, the risks of inaccuracy are significant. Poor analysis can lead to misguided investments, regulatory issues, and loss of trust. Existing financial benchmarks mainly evaluate how well LLMs answer financial questions but do not reflect performance in real-world tasks like generating financial analysis reports. In this paper, we propose FinAR-Bench, a solid benchmark dataset focusing on financial statement analysis, a core competence of fundamental analysis. To make the evaluation more precise and reliable, we break this task into three measurable steps: extracting key information, calculating financial indicators, and applying logical reasoning. This structured approach allows us to objectively assess how well LLMs perform each step of the process. Our findings offer a clear understanding of LLMs current strengths and limitations in fundamental analysis and provide a more practical way to benchmark their performance in real-world financial settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。