arXiv:2604.05912cs.CL2026-04被引 2

构建金融领域长期复杂任务基准,评估AI真实专业能力

FrontierFinance: A Long-Horizon Computer-Use Benchmark of Real-World Financial Tasks

  • 设计25个真实金融建模任务,每项需超18小时人工完成
  • 人类专家评分更高且更易产出可用结果,超越当前顶尖模型
  • 由金融从业者共建,覆盖行业标准流程与评分体系

随着人工智能在知识密集型领域引发就业替代担忧,现有评测基准无法衡量定义实际专业能力的任务表现。金融领域被认定为高风险受冲击领域,却缺乏追踪现实进展的可靠基准。当前大语言模型部署还缺乏明确责任机制,这一缺口亟待填补。为此,我们提出FrontierFinance,一个包含25个复杂金融建模任务的长期基准,覆盖五个核心金融模型,每项任务平均需超过18小时熟练人力完成。该基准由金融专业人士共同开发,反映行业标准建模流程,并配备详细评分标准以实现结构化评估。我们邀请人类专家参与任务定义、评分标准制定、模型评分及亲自执行任务作为人类基线。实验表明,人类专家在平均得分上更高,且更可能生成客户可用输出,优于当前最先进的系统。

原文摘要 · Abstract (English)

As concerns surrounding AI-driven labor displacement intensify in knowledge-intensive sectors, existing benchmarks fail to measure performance on tasks that define practical professional expertise. Finance, in particular, has been identified as a domain with high AI exposure risk, yet lacks robust benchmarks to track real-world developments. This gap is compounded by the absence of clear accountability mechanisms in current Large Language Model (LLM) deployments. To address this, we introduce FrontierFinance, a long-horizon benchmark of 25 complex financial modeling tasks across five core finance models, requiring an average of over 18 hours of skilled human labor per task to complete. Developed with financial professionals, the benchmark reflects industry-standard financial modeling workflows and is paired with detailed rubrics for structured evaluation. We engage human experts to define the tasks, create rubrics, grade LLMs, and perform the tasks themselves as human baselines. We demonstrate that our human experts both receive higher scores on average, and are more likely to provide client-ready outputs than current state-of-the-art systems.

金融AI评测基准长时任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。