arXiv:2604.21006cs.AIcs.LG2026-04被引 2

评测AI金融研究报告质量,发现仍不及人类专业水平。

Deep FinResearch Bench: Evaluating AI's Ability to Conduct Professional Financial Investment Research

论文配图:Deep FinResearch Bench: Evaluating AI's Ability to Conduct Professional Financial Investment Research
图 1 · 摘自论文原文
  • 构建三维度评估框架:定性严谨性、定量预测与估值、论点可信度。
  • 自动评分系统实现规模化评测,对比显示AI报告仍有差距。
  • 适合研究AI金融应用或评测基准的开发者参考。

我们提出 Deep FinResearch Bench,一个用于评估深度研究(DR)代理在金融投资研究中表现的实用且全面的评测框架。该基准从三个维度评估报告质量:定性严谨性、定量预测与估值准确性,以及论点的可信度和可验证性。特别地,我们定义了相应的定性与定量评估指标,并实现了自动化评分流程,支持规模化评估。将该框架应用于前沿深度研究代理生成的金融报告,并与金融专业人士撰写的报告进行对比,结果表明,当前AI生成报告在上述维度上仍存在明显不足。这些发现强调了针对金融领域定制化深度研究代理的必要性,我们希望本工作能为金融研究中深度研究代理的标准化评测奠定基础。

原文摘要 · Abstract (English)

We introduce Deep FinResearch Bench, a practical and comprehensive evaluation framework for deep research (DR) agents in financial investment research. The benchmark assesses three dimensions of report quality: qualitative rigor, quantitative forecasting and valuation accuracy, and claim credibility and verifiability. Particularly, we define corresponding qualitative and quantitative evaluation metrics and implement an automated scoring procedure to enable scalable assessment. Applying the benchmark to financial reports from frontier DR agents and comparing them with reports authored by financial professionals, we find that AI-generated reports still fall short across these dimensions. These findings underscore the need for domain-specialized DR agents tailored to finance, and we hope the work establishes a foundation for standardized benchmarking of DR agents in financial research.

金融AI评测基准深度研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。