arXiv:2507.16248cs.CL2025-07被引 13

用逻辑树评估金融研究智能体,自动判断其推理质量。

FinResearchBench: A Logic Tree based Agent-as-a-Judge Evaluation Framework for Financial Research Agents

  • 构建逻辑树作为中间结构,让智能体自评推理过程。
  • 覆盖70个典型金融研究问题,涵盖7类核心任务。
  • 适合评估金融领域复杂长程推理能力的AI系统。

近期,智能体在专业研究领域(如科学、软件开发和金融)迅速发展,其中深度研究智能体能完成长时序任务并解决复杂问题。然而,针对这类智能体的系统性、自动化评估框架仍十分匮乏。尤其在金融研究中,问题具有高度复杂性和细微差异。为此,我们提出FinResearchBench,一个基于逻辑树的“智能体即裁判”评估框架,专为金融研究智能体设计。该框架可自动评估智能体在金融研究领域7类关键任务中的表现,覆盖70个典型研究问题。主要贡献包括:(1)首个创新的“智能体即裁判”系统,通过提取研究结果的逻辑树作为中间信息,实现全面、可靠、稳健的评估;(2)面向金融领域,覆盖7种常见研究任务类型,具备高度针对性与实用性。

原文摘要 · Abstract (English)

Recently, AI agents are rapidly evolving in intelligence and widely used in professional research applications, such as STEM, software development, and finance. Among these AI agents, deep research agent is a key category as it can perform long-horizon tasks and solve problems of greater complexity. However, there are few evaluation frameworks and benchmarks that systematically and automatically investigate the capabilities of these research agents. In addition, financial research problems have distinct complexity and subtlety. To fill in the gap, we propose FinResearchBench, which is a logic tree-based Agent-as-a-Judge and targets specifically for the financial research agents. It provides a comprehensive and automatic assessment of the research agents across 7 key types of tasks in the financial research domain. The contributions of this work are two-folded: (1) the first and innovative Agent-as-a-Judge system that extracts the logic tree of the research outcome and uses it as the intermediate information to present a comprehensive, reliable, and robust evaluation; (2) finance-oriented that it covers 70 typical financial research questions, spreading across 7 frequently encountered types of task in the domain.

金融AI智能体评估逻辑树推理评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。