arXiv:2607.12252cs.CL2026-07被引 3

用大模型自动生成金融报告评价标准,实现大规模无专家评估。

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality

论文配图:FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality
图 1 · 摘自论文原文
  • 通过模型自动生成1.4万条评价标准,无需人工参与评分流程。
  • 98.67%的评价项与人工一致,确保大模型评分可替代人工。
  • 筛选出2600条有效标准,能清晰区分10个系统表现优劣。

深度研究代理在生成长篇金融报告方面日益普及,但大规模评估仍受限于需专家定义和执行高质量评价标准。本文提出一种无需人类专家参与最终环节的可扩展评价标准生成流程。基于104个真实用户查询,从模型生成的报告中自动合成14,450条特定查询的候选评价标准。为验证大模型评价的有效性,我们在抽样数据上对比三位人类专家与三模型评委的判断,结果显示两者在联合一致项上达到98.67%的标签级一致性。随后通过双重过滤:严格一致性过滤(要求三模型评委对同一查询下所有报告均一致)与可区分性过滤(要求标准至少在一项上给出多数“是”和多数“否”),保留3,687条一致性通过的标准,其中2,600条具有可区分性,构成最终共识黄金标准集。利用该标准集,对10个深度研究系统进行评估,项目通过率范围为58.58%至22.23%,实现了清晰的性能区分。该流程完全摆脱人工专家执行环节,天然适用于大规模基准测试、系统自动对比及以评估驱动的系统优化研究。

原文摘要 · Abstract (English)

Deep research agents are increasingly used to produce long-form financial reports, yet large-scale evaluation remains bottlenecked by the need for human experts to define and execute high-quality rubrics. We address this problem by proposing a scalable pipeline for generating high-quality rubrics without human experts in the final loop. We build a financial deep research benchmark from 104 real-world user queries and automatically synthesize 14,450 query-specific candidate rubrics from model-generated reports. To justify removing human experts from rubric execution, we compare rubric judgments from three human experts with those from a three-LLM judge panel on a sampled subset, and show that LLM-based evaluation is sufficiently consistent with human evaluation to replace it for large-scale rubric screening, including 98.67\% label-level agreement on jointly unanimous items. We then derive consensus-derived gold rubrics through two filters: a strict consistency filter, which keeps a rubric only if the three LLM judges unanimously agree on every report under the same query, and a distinguishability filter, which keeps a rubric only if it assigns at least one majority-yes and at least one majority-no label across the evaluated systems. This process retains 3,687 consistency-passed rubrics, of which 2,600 remain distinguishable and form the final set of consensus-derived gold rubrics. Using this final rubric set, we obtain clearly differentiated rankings across 10 deep research systems, with item-level pass rates ranging from 58.58\% to 22.23\%. More broadly, because the pipeline removes human-expert execution from rubric generation and evaluation, it is naturally scalable for benchmark evaluation, automatic system comparison, and future studies of evaluation-driven system improvement.

金融分析评估基准大模型评测自动化评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。