arXiv:2608.18289cs.AIcs.CR2026-08中稿 · Workshop on System…

评测开源模型在高风险公共领域文档提取中的表现,发现多数配置效果不佳。

Evaluating Structured Information Extraction with Open Models in a High Risk Public Sector Application

论文配图:Evaluating Structured Information Extraction with Open Models in a High Risk Public Sector Application
图 1 · 摘自论文原文
  • 构建真实学生申请任务的端到端评估基准,测试开源OCR+LLM/VLM组合
  • 35种配置中仅4种F1超0.5,多数低于0.25,大模型不保证更好结果
  • OCR输出结构保留度比下游模型更重要,适合公共安全场景研究者参考

从非结构化文档中提取结构化信息是各行业数字化转型的关键环节。尽管商业应用多依赖专有方案,但开源的OCR引擎、大语言模型(LLMs)和视觉语言模型(VLMs)正快速发展,提供可访问替代方案。然而,针对真实、多步骤提取流程的系统性评估仍十分匮乏。在欧盟人工智能法案将此类应用划为高风险的背景下,负责任地使用这些工具需在真实任务上进行全面评估。为此,本文提出一个综合性基准,评估开源系统在一项被归类为高风险的真实世界文档处理任务——国际学习项目学生申请中的端到端性能。我们对最先进的OCR引擎、LLMs和VLMs进行了全面实证评估。结果显示,尽管VLM整体优于OCR+LLM流水线,但即使是先进的开源模型,在零样本设置下也难以可靠完成任务。35种配置中仅有4种的F1得分高于0.5,最佳的OCR+LLM流水线与顶级VLM性能相当,但大多数组合表现显著更差。约75%的配置得分低于0.25。模型规模影响性能,但关系非线性:模型越大并不意味着成比例提升。输入质量,尤其是OCR输出的结构保真度,成为独立于下游模型能力的关键因素。

原文摘要 · Abstract (English)

The extraction of structured information from unstructured documents represents a critical component of digital transformations in all sectors. While proprietary solutions dominate commercial applications, a rapidly growing ecosystem of open-source Optical Character Recognition (OCR) engines, Large Language Models (LLMs), and Vision-Language Models (VLMs) offers accessible alternatives. However, systematic evaluations on realistic, multi-step extraction pipelines remain scarce. Responsible usage of such extraction tools require comprehensive evaluations on realistic tasks, especially as these solutions will be key components of applications in the public sector that the EU AI act categorizes as high risk. To address this gap we present a comprehensive benchmark assessing the end-to-end performance of open-source systems on a complex real-world document processing task classified as high risk: Student applications for an international study program. We conduct a comprehensive empirical evaluation with state-of-the-art OCR engines, LLMs and VLMs. Our results reveal that while VLMs generally outperform OCR+LLM pipelines, even state-of-the-art open-source models struggle to handle such tasks reliably in zero-shot settings. Only 4 of 35 configurations achieved F1 scores above 0.5, with the best OCR+LLM pipeline matching top VLM performance, though most OCR+LLM combinations performed substantially worse. Roughly 75\% of all configurations scored below 0.25. Model scale influences performance, yet the relationship is non-linear: substantially larger models do not guarantee proportionally better results. Input quality, particularly the structural preservation of OCR output, emerges as a critical factor independent of downstream model capability.

信息抽取开源模型高风险应用文档理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。