arXiv:2606.03829cs.AI2026-06被引 4

构建可审计的金融研究评估基准,聚焦推导过程而非仅答案。

BigFinanceBench: A Workflow-Grounded Benchmark for Financial-Research Agents

论文配图:BigFinanceBench: A Workflow-Grounded Benchmark for Financial-Research Agents
图 1 · 摘自论文原文
  • 以可拆解步骤的评分规则评估完整分析流程
  • 928个任务覆盖真实金融研究场景,支持部分得分与错误定位
  • 发现当前模型推导能力普遍不足,答案准确率不能反映过程质量

金融研究结论只有在可被同行审计时才具决策价值:包括选用的来源、时间范围与会计定义、假设设定及计算方式。现有金融评测多关注孤立子技能或最终答案,对可审计推导过程关注不足。我们提出 BigFinanceBench,一个由专家撰写的928项开放式金融研究任务基准,每个任务配有一份基于分步权重的评分标准,将推导过程分解为可独立核查的步骤。该基准以工作流为基础,评估完整推导链条而非仅最终输出。在36,241个评分点上,支持部分得分与流程中故障的精准定位。对十种前沿及开源智能体的评估显示,存在显著提升空间:最优系统仅获58.8%评分;最终答案准确率是推导质量的有损代理;模型能力在不同金融工作流中表现不均。

原文摘要 · Abstract (English)

Financial-research answers are decision-relevant only when another analyst can audit how they were produced: which source was chosen, which period and accounting definition were used, which assumptions were made, and how the calculation was performed. Existing finance benchmarks largely evaluate isolated subskills or final answers, leaving the auditable derivation itself under-measured. We introduce BigFinanceBench, a 928-item expert-authored benchmark of open-ended financial-research tasks in which each item pairs a ground-truth reference answer with a point-weighted rubric that decomposes the derivation into independently checkable steps. BigFinanceBench is workflow-grounded in that it evaluates the full derivation rather than only the final output. Across 36,241 rubric points, the benchmark supports partial-credit evaluation and localization of failures across the analyst workflow. Evaluating ten current frontier and open-weight agents, we find substantial headroom: the best system reaches only 58.8% rubric score, final-answer accuracy is a useful but lossy proxy for derivation quality, and model capability varies non-uniformly across financial workflows.

金融AI评估基准可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。