arXiv:2605.14355cs.AIcs.CL2026-05被引 3

首个评估金融智能体全流程能力的基准,揭示当前模型在复杂任务中的短板。

Herculean: An Agentic Benchmark for Financial Intelligence

论文配图:Herculean: An Agentic Benchmark for Financial Intelligence
图 1 · 摘自论文原文
  • 构建四类金融工作流的标准化智能体环境,支持端到端评估
  • 前沿模型在交易与洞察任务表现良好,但在对冲与审计中严重受限
  • 凸显长期协调与状态一致性对高风险金融任务的关键作用

随着人工智能代理的进步,核心问题已不再是能否完成孤立的金融任务,而是能否可靠执行金融专业工作。现有金融基准主要评估问答、检索、摘要和分类等静态能力,难以全面反映实际应用水平。我们提出Herculean,首个面向代理型金融智能的综合性基准,涵盖交易、对冲、市场洞察和审计四类典型工作流。每个工作流均基于MCP框架构建标准化技能环境,包含专属工具、交互机制、约束条件和成功标准,实现对异构代理系统的统一端到端评估。实验发现,前沿代理在交易与市场洞察任务上表现较好,但在对冲与审计任务中显著受阻,原因在于长期协调、状态一致性和结构化验证的重要性。整体结果表明,当前智能体在将金融推理转化为高风险场景下的可靠工作流执行方面仍存在关键缺口。

原文摘要 · Abstract (English)

As AI agents improve, the central question is no longer whether they can solve isolated well-defined financial tasks, but whether they can reliably carry out financial professional work. Existing financial benchmarks offer only a partial view of this ability, as they primarily evaluate static competencies such as question answering, retrieval, summarization, and classification. We introduce Herculean, the first skilled benchmark for agentic financial intelligence spanning four representative workflows, including Trading, Hedging, Market Insights, and Auditing. Each workflow is instantiated as a standardized MCP-based skill environment with its own tools, interaction dynamics, constraints, and success criteria, enabling consistent end-to-end assessment of heterogeneous agent systems. Across frontier agents, we find agents perform relatively well on Trading and Market Insights, but struggle substantially on Hedging and Auditing, where long-horizon coordination, state consistency, and structured verification are critical. Overall, our results point to a key gap in current agents in turning financial reasoning into dependable workflow execution in high-stakes financial workflows.

金融智能智能体评测工作流评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。