构建首个覆盖完整投资流程的金融智能评估基准,揭示工具链比模型本身更关键。
FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents

- 设计220个专家级问题与1.15万条带来源标注的评分标准,覆盖全投资流程
- 最强模型(Claude Fable 5)仅达49.2%,最优系统(Samaya)以2.2倍成本优势领先
- 开源数据集和评测代码,适合研究金融AI代理与评估体系的学者
AI代理在专业投资研究中日益应用,但现有基准仅聚焦金融数据提取这一狭窄领域,且当前模型已接近饱和。参考文献评分和通用LLM作为评判者的方法难以应对分析师所需开放式、长文本回答的真实需求。我们提出FrontierFinance,一个完全开放的基准,包含220个专家设计的问题和11,543条带来源标注的评分标准,覆盖投资全流程中的六大关键应用场景。该基准比现有公开金融基准更广泛也更困难。在仅使用公开数据的统一测试环境下评估前沿模型与代理系统,发现工具链对质量与效率影响远超模型本身;自研系统Samaya以56.0%的得分领先最强前沿模型Claude Fable 5(49.2%),成本低约2.2倍;最佳开源模型Kimi K3(46.4%)接近最优专有模型,成本仅为其1/4.5。筛查与发现、行业与宏观分析仍是所有系统最难任务,最高得分分别为33%和39%。数据集与评分代码已公开。
原文摘要 · Abstract (English)
AI agents are increasingly deployed for professional investment research, yet no benchmark captures the complexity of the full investor workflow. Existing benchmarks mainly target financial data extraction, a narrow slice that current models have largely saturated, while reference-based metrics and generic LLM-as-a-judge scoring fall short on the open-ended, long-form answers that real analyst queries demand. We introduce FrontierFinance, a fully open benchmark of 220 expert-crafted queries and 11,543 source-attributed rubrics spanning six crucial use cases across the full investor workflow. FrontierFinance is both broader and harder than existing public finance benchmarks. Evaluating frontier models and agent systems under a common harness restricted to publicly available data, we find that the tool harness, not the model alone, strongly shapes quality and efficiency; that Samaya's in-house system leads at 56.0%, ahead of the strongest frontier model (Claude Fable 5, 49.2%) at roughly 2.2x lower cost; and that the best open-weight model (Kimi K3, 46.4%) nearly matches the best proprietary model at 4.5x lower cost. Screening & Discovery and Sector, Industry & Macro remain the hardest use cases across all systems, where even the best systems reach only 33% and 39%. We make the dataset and grading code publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。