用专业人士成果构建评分标准,让金融AI评估更真实可靠。
FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables

- 从从业者实际交付物中提取隐性专业标准,构建可复用评分体系。
- 在专业角色特殊任务上,新方法比仅靠提示词提升21.1个百分点。
- 适合金融AI研发、评测人员及需要高精度评估的机构使用。
评估金融AI代理需匹配真实职业标准。现有评分方法多基于任务提示或模型输出,忽略仅在从业者交付物中显现的隐性规范。本文提出FinProBench基准与角色根基评分构建(RGRC)流程,从同一角色从业者交付物中提取评分标准。RGRC包含四阶段:交付物收集、能力提取、评分合成与验证。其标准捕捉隐性规范,区分质量层级,并可在角色内跨任务复用。我们先将57种职业按交付物类型划分为30类具丰富先验的常规角色和27类先验稀疏的角色专精角色。在所有角色中,仅用提示词的方法在常规角色上接近RGRC(89.2% vs. 90.7%),但在角色专精任务上显著落后(78.0% vs. 99.1%)。这表明当模型先验充分时提示工程可近似评分,而超出先验的规范必须依赖专业根基。FinProBench涵盖1,723份精心筛选的交付物,覆盖57种职业、8个金融子行业和161种交付物类型,发布初始20个完整任务评估集,覆盖20个角色和7个子行业。采用异构LLM裁判和角色级评分标准,人类交付物平均得分73.7(满分100),高于四个系统(70.3、70.2、69.6),各系统置信区间重叠且互补。角色级复用评分可使每任务构建成本降低6.7倍。
原文摘要 · Abstract (English)
Evaluating financial AI agents requires criteria aligned with real professional work. Existing rubric methods typically derive criteria from task prompts or model outputs, overlooking tacit standards visible only in practitioner deliverables. We introduce FinProBench, a benchmark for professional financial tasks, and Role-Grounded Rubric Construction (RGRC), a reusable pipeline that derives rubrics from deliverables produced by practitioners in the same role. RGRC comprises four stages: Deliverable Collection, Competency Extraction, Rubric Synthesis, and Validation. Its rubrics capture tacit standards, distinguish quality levels, and transfer across tasks within a role. Before analysis, we classified 57 occupations by deliverable genre into 30 prior-rich conventional roles and 27 prior-sparse role-specialized roles. Across all roles, Prompt-only nearly matches RGRC for conventional roles (89.2% vs. 90.7%), but RGRC substantially outperforms it for role-specialized roles (99.1% vs. 78.0%). This split indicates that prompt engineering can approximate rubrics when conventions are well represented in model priors, while professional grounding is essential for standards beyond those priors. FinProBench is built from 1,723 curated deliverables spanning 57 occupations, 8 financial sub-industries, and 161 deliverable types, and releases an initial evaluation set of 20 complete tasks covering 20 roles in 7 sub-industries. With heterogeneous LLM judges and role-level rubrics, human deliverables rank first on average (73.7 vs. 70.3, 70.2, and 69.6 out of 100), while all four systems show overlapping 95% confidence intervals and complementary strengths. Reusing rubrics at the role level reduces estimated per-task construction effort by 6.7 times relative to authoring each rubric from scratch.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。