arXiv:2608.00764cs.AI2026-08

首个评估金融指标生成全流程的基准,揭示大模型在数据获取环节严重失准。

FinDeepIndicator: Benchmarking Deep Research Agents in End-to-End Financial Indicator Construction

论文配图:FinDeepIndicator: Benchmarking Deep Research Agents in End-to-End Financial Indicator Construction
图 1 · 摘自论文原文
  • 构建四阶段全流程评测框架:公式定义→数据收集→计算→答案生成
  • 覆盖3350个问题、800家上市公司、10年历史数据,含中美市场样本
  • 大模型公式准确但数据检索与计算失误多,专业代理仍不可靠

财务指标是将原始财务数据转化为可解释度量的关键工具,广泛应用于估值、风险评估和经济分析等下游任务。然而,现有金融评测主要关注答案准确性,且假定相关数据已提供,对指标构建过程中的中间步骤缺乏评估。本文提出FinDeepIndicator,首个专注于端到端财务指标构建中深度研究(DR)代理的评测基准。该基准涵盖公式定义、数据收集、指标计算和答案生成四个阶段,包含21个细粒度子类别的基础、技术及宏观经济指标。数据集包含3,350个精心筛选的问答对,源自美中两个市场的数据,涵盖10年历史财务数据及800家上市公司。对配备搜索功能的大语言模型(LLMs)和DR代理的大量实验表明:尽管LLMs在公式定义上表现良好,但在数据检索和数值执行阶段准确率显著下降;而DR代理整体优于搜索型LLMs,但在真实金融分析场景中仍不可靠。这些发现为开发更强大、可信的金融领域研究代理提供了重要启示。

原文摘要 · Abstract (English)

Financial indicators are essential tools for transforming raw financial data into interpretable measures for various downstream tasks, such as valuation, risk assessment, and economic analysis. However, existing financial benchmarks largely focus on answer-level accuracy and often assume that relevant data are already provided, leaving the assessment of the intermediate process of indicator construction underexplored. In this work, we propose FinDeepIndicator, the first benchmark dedicated to evaluating Deep Research (DR) agents in end-to-end financial indicator construction. Specifically, FinDeepIndicator evaluates DR agents across four stages in indicator construction: formula specification, data collection, indicator calculation, and answer generation, and covers fundamental, technical, and macroeconomic indicators organized into 21 fine-grained sub-categories. It contains 3,350 curated question-answer (QA) pairs derived from both U.S. and Chinese markets, 10 years of historical financial data, and 800 listed companies. Extensive experiments on search-equipped Large Language Models (LLMs) and DR agents show that, while LLMs generally perform well in formula specification, their accuracy drops substantially during data retrieval and numerical execution. DR agents consistently outperform search-equipped LLMs, yet remain unreliable in realistic financial analysis settings. These findings provide insights for developing more capable and trustworthy DR agents in finance.

金融AI评测基准研究代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。