arXiv:2511.14998cs.CV2025-11

评测金融文档OCR是否准确保留关键事实,发现数值和货币单位最易出错。

FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR

  • 用规则引导的LLM评判模型输出是否保留事实,而非仅比对文字
  • 859页真实财务文件中9481个关键事实被标注,数值类错误最多
  • 适合关注高风险场景下文档理解可靠性的研究者和开发者

多模态大模型虽提升文档理解能力,但仅在表面指标上表现好,并不意味着能忠实保留决策关键证据。尤其在金融文档中,微小视觉误差可能导致语义显著变化。为此,我们提出FinCriticalED(金融关键错误检测),一个以事实为中心的视觉基准,用于评估OCR与视觉语言系统是否在上下文中保留金融关键信息。该基准包含859张真实金融文档页面,涵盖9481个专家标注的事实,涉及五类关键字段:数值、时间、货币单位、报告主体和金融概念。我们将任务定义为结构化OCR与事实级验证,并设计基于确定性规则的LLM作为裁判协议,判断模型输出是否保持原标注事实。我们对13种系统进行评测,包括传统OCR流程、专用OCR-VLM、开源及专有MLLM。结果显示,字面准确率与事实可靠性之间存在明显差距,其中数值和货币单位最为脆弱,且错误集中于布局复杂的混合型文档,不同模型家族呈现不同失效模式。总体而言,FinCriticalED为可信金融OCR提供了严格基准,也为高风险多模态文档理解中的证据保真度提供实用测试平台。数据集与基准详情见https://the-finai.github.io/FinCriticalED/

原文摘要 · Abstract (English)

Recent progress in multimodal large language models (MLLMs) has substantially improved document understanding, yet strong optical character recognition (OCR) performance on surface metrics does not guarantee faithful preservation of decision-critical evidence. This limitation is especially consequential in financial documents, where small visual errors can induce discrete shifts in meaning. To study this gap, we introduce FinCriticalED (Financial Critical Error Detection), a fact-centric visual benchmark for evaluating whether OCR and vision-language systems preserve financially critical evidence beyond lexical similarity. FinCriticalED contains 859 real-world financial document pages with 9,481 expert-annotated facts spanning five critical field types: numeric, temporal, monetary unit, reporting entity, and financial concept. We formulate the task as structured OCR with fact-level verification, and develop a Deterministic-Rule-Guided LLM-as-Judge protocol to assess whether model outputs preserve annotated facts in context. We benchmark 13 systems spanning OCR pipelines, specialized OCR VLMs, open-source MLLMs, and proprietary MLLMs. Results reveal a clear gap between lexical accuracy and factual reliability, with numerical values and monetary units emerging as the most vulnerable fact types, and critical errors concentrating in visually complex, mixed-layout documents with distinct failure patterns across model families. Overall, FinCriticalED provides a rigorous benchmark for trustworthy financial OCR and a practical testbed for evidence fidelity in high-stakes multimodal document understanding. Benchmark and dataset details available at https://the-finai.github.io/FinCriticalED/

OCR金融文档事实验证多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。