arXiv:2510.13936cs.CL2025-10被引 11

首个系统评估金融研究智能体的基准,覆盖多语言多市场

FinDeepResearch: Evaluating Deep Research Agents in Rigorous Financial Analysis

  • 构建分层评分框架,模拟专业分析师从数据到策略的全流程
  • 涵盖64家上市公司、15808个评分项,覆盖4语言8市场
  • 对比16种方法,揭示不同模型在多场景下的优劣表现

深度研究(DR)智能体依托先进大语言模型,在复杂研究任务中展现出强大潜力。然而,现有研究缺乏对DR智能体在关键研究分析中能力的严谨系统评估。为此,我们提出HisRubric,一种具有分层分析结构和细粒度评分标准的新评估框架,严格评估DR智能体在企业财务分析中的能力。该框架模拟专业分析师的工作流程,从数据识别、指标计算,到战略总结与解读逐步推进。基于此框架,我们构建了FinDeepResearch基准,包含来自4个语言、8个金融市场的64家上市公司,共计15,808个评分项。我们在该基准上对16种代表性方法进行了广泛实验,包括6个DR智能体、5个具备深度推理与搜索能力的LLM,以及5个仅具深度推理能力的LLM。结果揭示了各类方法在不同能力维度、金融市场及语言环境中的优势与局限,为未来研究与开发提供了宝贵洞见。基准与评估代码已公开于https://OpenFinArena.com/。

原文摘要 · Abstract (English)

Deep Research (DR) agents, powered by advanced Large Language Models (LLMs), have recently garnered increasing attention for their capability in conducting complex research tasks. However, existing literature lacks a rigorous and systematic evaluation of DR Agent's capabilities in critical research analysis. To address this gap, we first propose HisRubric, a novel evaluation framework with a hierarchical analytical structure and a fine-grained grading rubric for rigorously assessing DR agents' capabilities in corporate financial analysis. This framework mirrors the professional analyst's workflow, progressing from data recognition to metric calculation, and finally to strategic summarization and interpretation. Built on this framework, we construct a FinDeepResearch benchmark that comprises 64 listed companies from 8 financial markets across 4 languages, encompassing a total of 15,808 grading items. We further conduct extensive experiments on the FinDeepResearch using 16 representative methods, including 6 DR agents, 5 LLMs equipped with both deep reasoning and search capabilities, and 5 LLMs with deep reasoning capabilities only. The results reveal the strengths and limitations of these approaches across diverse capabilities, financial markets, and languages, offering valuable insights for future research and development. The benchmark and evaluation code is publicly available at https://OpenFinArena.com/.

金融分析智能体评估多语言基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。