多语言表格提取评测基准,含结构化评估方法。
PulseBench-Tab: A Multilingual Benchmark for Table Extraction with Graph-Based Evaluation

- 构建跨9语言4文字的1820张真实文档表格数据集。
- 48.1%表格含合并单元格,最大达1183个单元格。
- 提出图结构匹配评分法,精准衡量提取结果质量。
我们提出PulseBench-Tab,一个开放的多语言表格提取评测基准。该基准包含1820张由人工标注的表格,覆盖9种语言和4种书写系统(拉丁、中文日文韩文、阿拉伯、西里尔),数据源涵盖财务报告、政府文件和监管披露等380份真实文档。表格规模从2到1183个单元格不等,其中48.1%包含合并或跨行单元格。我们同时提出T-LAG(表逻辑邻接图)评估指标,将表格建模为基于单元格邻接关系的有向图,通过最优二分匹配在单一分数中同时衡量结构与内容保真度。我们在该基准上评估了9个商用及开源表格提取系统,并提供各语言细分结果。完整数据集、评分代码及所有系统输出均已公开。
原文摘要 · Abstract (English)
We introduce PulseBench-Tab, an open multilingual benchmark for evaluating table extraction from document images. The benchmark comprises 1,820 human-annotated tables spanning 9 languages and 4 scripts (Latin, CJK, Arabic, Cyrillic), drawn from 380 real-world source documents including financial filings, government reports, and regulatory disclosures. Tables range from 2 to 1,183 cells, with 48.1% containing merged or spanning cells. Alongside the dataset, we propose T-LAG (Table Logical Adjacency Graph), a novel evaluation metric that models tables as directed graphs over cell adjacencies and computes structural and content fidelity in a single score via optimal bipartite matching. We evaluate 9 commercial and open-source table extraction systems across the benchmark and report per-language breakdowns. The full dataset, scoring code, and all provider outputs are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。