arXiv:2605.29586cs.AI2026-05被引 1

评测大模型财务报表一致性验证能力,发现其误报率高达100%。

FinVerBench: Benchmark Validity and Calibration in Large Language Model Financial Statement Verification

论文配图:FinVerBench: Benchmark Validity and Calibration in Large Language Model Financial Statement Verification
图 1 · 摘自论文原文
  • 构建43家标普500公司财报数据集,定义四类错误类型
  • 14个大模型中9个在干净报表上误报率达95%-100%
  • 结果受数据呈现方式影响,需考虑真实财务显示环境

我们提出FinVerBench,一个用于财务报表一致性验证的基准测试与有效性研究:判断企业财报中的数值是否在模型可见信息范围内逻辑一致。该基准基于43家标普500公司从SEC 10-K XBRL文件中提取的数据,定义了四类错误分类:算术错误、跨表链接错误、年同比错误和数值扰动。我们对15个主流大模型进行了评估,报告了14次完整运行;其中一次Gemini 2.5 Pro因40/108个接口调用失败被排除主比较。所有二分类指标均剔除未确定为正例的实例(因扰动项目未显示),保留105个可观察诊断样本(43个无错,62个注入错误)。在原始引导式清单提示下,14次完整运行中有9次在干净报表上产生95%-100%假阳性,仅1次实现0%观测假阳性率。基准渲染方式显著影响测量召回率:在真实场景的舍入版本上,校准模型召回率为79.0%,假阳性率为0%;而在未舍入诊断变体上,召回率可达100.0%。结果支持构建效度结论而非最终排行榜:财务报表验证不仅是算术检测,更涉及不完全可观测性下的校准判断、提示诱导假设与真实数值呈现。FinVerBench及全部代码已公开。

原文摘要 · Abstract (English)

We introduce FinVerBench, a benchmark and validity study for financial statement verification: determining whether a set of corporate financial statements is numerically consistent from the information shown to the model. FinVerBench is built from SEC 10-K XBRL filings for 43 S&P 500 companies and defines a four-category error taxonomy covering arithmetic, cross-statement linkage, year-over-year, and magnitude perturbations. We attempt fifteen contemporary LLM evaluations and report fourteen complete runs; a Gemini 2.5 Pro run is excluded from the main comparison because 40/108 gateway calls failed. All binary metrics exclude underdetermined positive instances whose perturbed line item is not rendered, leaving a 105-instance observable diagnostic subset (43 clean, 62 error-injected). Under the original guided-checklist prompt on the unrounded diagnostic subset, nine of fourteen complete LLM runs produce 95-100% false positives on clean statements, while one run achieves 0% observed false positives. Benchmark rendering choices materially affect measured recall: on a realistic rounded variant of the same observable subset, the calibrated model's recall is 79.0% with 0% observed FPR, compared with 100.0% recall on the unrounded diagnostic variant. These results support a construct-validity conclusion rather than a final leaderboard: financial statement verification is not merely arithmetic detection, but calibrated judgment under incomplete observability, prompt-induced assumptions, and realistic numerical rendering. FinVerBench and all code are publicly available.

财务验证大模型评测基准测试算术推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。