arXiv:2509.19345cs.CLcs.AI2025-09被引 6

为生成式文档解析设计公平评估框架,解决传统指标误判合理差异的问题。

SCORE: A Semantic Evaluation Framework for Generative Document Parsing

  • 提出无解释依赖的评分框架,融合内容鲁棒性与结构一致性评估
  • 在1114页数据上发现传统指标平均误判12%-25%且扭曲排名
  • 支持多维度诊断,适合研究生成式文档解析的开发者和评测者

多模态生成式文档解析系统挑战传统评估方式:与确定性OCR或布局模型不同,它们常产生语义正确但结构不同的输出。传统指标(如CER、WER、IoU或TEDS)将此类多样性误判为错误,惩罚有效解读并掩盖系统行为。我们提出SCORE(结构与内容鲁棒评估),一个不依赖解释的评估框架,包含:(i) 改进编辑距离以实现内容保真度鲁棒性,(ii) 词级诊断区分幻觉与遗漏,(iii) 带空间容差与语义对齐的表格评估,(iv) 层次感知的一致性检查。这些维度共同实现对表示多样性包容的同时保持语义严谨。在涵盖1,114页的综合基准和真实场景数据集上,SCORE持续揭示了传统指标忽略的跨数据集性能模式。在表结构模糊的2%-5%页面中,传统指标平均使系统得分下降12%-25%,导致排名失真。SCORE修正了这些情况,恢复了不同但有效的解释间的等价性。此外,通过将生成输出标准化为格式无关表示,SCORE在无需目标检测流水线的情况下复现了传统分数(如表格F1高达0.93),证明仅靠生成式解析即可完成全面评估。通过揭示解释多样性对评估结果的影响,并提供多维度可解释诊断,SCORE确立了现代文档解析系统语义基础、公平且实用的评测原则。

原文摘要 · Abstract (English)

Multi-modal generative document parsing systems challenge traditional evaluation: unlike deterministic OCR or layout models, they often produce semantically correct yet structurally divergent outputs. Conventional metrics-CER, WER, IoU, or TEDS-misclassify such diversity as error, penalizing valid interpretations and obscuring system behavior. We introduce SCORE (Structural and COntent Robust Evaluation), an interpretation-agnostic framework that integrates (i) adjusted edit distance for robust content fidelity, (ii) token-level diagnostics to distinguish hallucinations from omissions, (iii) table evaluation with spatial tolerance and semantic alignment, and (iv) hierarchy-aware consistency checks. Together, these dimensions enable evaluation that embraces representational diversity while enforcing semantic rigor. Across 1,114 pages spanning a holistic benchmark and a field dataset, SCORE consistently revealed cross-dataset performance patterns missed by standard metrics. In 2-5% of pages with ambiguous table structures, traditional metrics penalized systems by 12-25% on average, leading to distorted rankings. SCORE corrected these cases, recovering equivalence between alternative but valid interpretations. Moreover, by normalizing generative outputs into a format-agnostic representation, SCORE reproduces traditional scores (e.g., table F1 up to 0.93) without requiring object-detection pipelines, demonstrating that generative parsing alone suffices for comprehensive evaluation. By exposing how interpretive diversity impacts evaluation outcomes and providing multi-dimensional, interpretable diagnostics, SCORE establishes foundational principles for semantically grounded, fair, and practical benchmarking of modern document parsing systems.

文档解析生成评估语义评价多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。