arXiv:2604.11304cs.AI2026-04被引 2

评测AI在投行全流程中的表现,发现顶尖模型仍难胜任真实工作。

BankerToolBench: Evaluating AI Agents in End-to-End Investment Banking Workflows

  • 构建真实投行工作流的开源评测基准,需跨工具协同生成多文件成果。
  • 顶级模型仅通过约一半评估标准,0%输出被银行家认为可交付客户。
  • 揭示跨文档一致性等关键缺陷,为高价值职业AI应用提供改进方向。

现有AI评测缺乏对专业工作流中经济意义进展的衡量能力。为评估前沿AI代理在高价值、劳动密集型职业中的表现,我们提出BankerToolBench(BTB):一个面向初级投行分析师日常任务的端到端分析工作流开源基准。为确保生态有效性,我们与502名来自头部投行的从业者合作开发。BTB要求代理通过导航数据室、使用行业工具(如市场数据平台、SEC文件数据库)并生成包含Excel财务模型、演示文稿及PDF/Word报告的多文件成果来响应资深投行人员请求。完成一项BTB任务耗时高达21小时,凸显将此类工作委托给AI的经济重要性。BTB支持对任意大模型或代理进行自动化评估,其评分基于超过100项由经验丰富的投行人员定义的评分标准,以捕捉利益相关方的实际效用。测试9个前沿模型后发现,即使表现最优的GPT-5.4也未能通过近半数标准,且银行家对其输出的客户可用性评分为0%。失败分析揭示了跨成果一致性等关键障碍,并指明了代理式AI在高风险专业工作流中的改进路径。

原文摘要 · Abstract (English)

Existing AI benchmarks lack the fidelity to assess economically meaningful progress on professional workflows. To evaluate frontier AI agents in a high-value, labor-intensive profession, we introduce BankerToolBench (BTB): an open-source benchmark of end-to-end analytical workflows routinely performed by junior investment bankers. To develop an ecologically valid benchmark grounded in representative work environments, we collaborated with 502 investment bankers from leading firms. BTB requires agents to execute senior banker requests by navigating data rooms, using industry tools (market data platform, SEC filings database), and generating multi-file deliverables--including Excel financial models, PowerPoint pitch decks, and PDF/Word reports. Completing a BTB task takes bankers up to 21 hours, underscoring the economic stakes of successfully delegating this work to AI. BTB enables automated evaluation of any LLM or agent, scoring deliverables against 100+ rubric criteria defined by veteran investment bankers to capture stakeholder utility. Testing 9 frontier models, we find that even the best-performing model (GPT-5.4) fails nearly half of the rubric criteria and bankers rate 0% of its outputs as client-ready. Our failure analysis reveals key obstacles (such as breakdowns in cross-artifact consistency) and improvement directions for agentic AI in high-stakes professional workflows.

AI代理投行评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。