arXiv:2608.04374cs.CLcs.AI2026-08

评测并提升大模型生成机构级财务报告的能力,解决内容可信度与专业性问题。

FinReportBench: Measuring and Improving Institution-Grade Financial Report Generation

论文配图:FinReportBench: Measuring and Improving Institution-Grade Financial Report Generation
图 1 · 摘自论文原文
  • 基于专家评审构建35项评分标准,涵盖报告可交付性、身份一致性和机构完整性。
  • 在244个中英文双语任务上测试,发现报告身份和机构完整性仍是主要短板。
  • 通过技能蒸馏改进模型,使生成质量平均提升33.85分,适合金融研究者使用。

大语言模型能生成流畅的财务分析,但流畅不等于适合机构使用。本文提出FinReportBench,一个基于专家评审的基准,用于衡量和改进机构级财务报告生成。专家评审揭示报告身份、机构组件、来源严谨性及视觉呈现等方面的普遍缺失。我们通过专家部分排序、多模态证据和决策边界审计,构建了包含35项指标的评分体系,覆盖可交付性、报告身份和机构完整性。基于10,000条平衡的中英文金融研究源数据,构建了244个双语任务,涵盖三个研究对象和两层输入。每个任务分离公开查询、重构的研究轨迹和隐藏的源数据包。三组独立评审员以接近天花板的准确率复现了专家排序,表明有限且可观测的标准可支持可靠评估。在九个模型家族中,基础可交付性几乎饱和,而报告身份和机构完整性仍是主要瓶颈。最大跨模型差距集中在生成追踪控制、信息密度和数据纪律,而非基本报告框架。随后,我们采用基准引导的技能蒸馏,将重复失败转化为可复用的生成与自检约束。在五个模型家族中,优化后的技能使平均G1提升33.85分,平均G2提升13.83分,同时每对均保持G0水平。代码与基准资源见https://github.com/MisterBrookT/finreportbench。

原文摘要 · Abstract (English)

Large language models can produce fluent financial analysis, but fluency alone does not establish whether a report is suitable for institutional delivery. We introduce FinReportBench, an expert-grounded benchmark for measuring and improving institution-grade financial report generation. Expert review reveals recurring gaps in report identity, institutional components, source discipline, and visual delivery. We derive a 35-item rubric through expert partial orders, multimodal evidence, and audits of decision boundaries, covering deliverability, report identity, and institutional completeness. Starting from 10,000 balanced Chinese and English financial-research source records, we curate 244 bilingual tasks across three research objects and two input tiers. Each task separates the public query, reconstructed research trajectory, and hidden source packet. Three independent judge families reproduce the expert partial order at near-ceiling rates, showing that bounded, observable criteria support reliable evaluation. Across nine model families, basic deliverability is nearly saturated, while report identity and institutional completeness remain the primary bottlenecks. The largest cross-model gaps concern generation-trace control, information density, and data discipline rather than basic report framing. We then use benchmark-guided skill distillation to turn recurrent failures into reusable generation and self-review constraints. Across five model families, the evolved skill improves mean G1 by 33.85 points and mean G2 by 13.83 points over paired no-skill runs while preserving G0 for every pair. Code and benchmark artifacts are available at https://github.com/MisterBrookT/finreportbench.

财务报告大模型评估技能蒸馏多模态评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。