让表格与上下文语义对齐,实现文档深层理解。
From Surface to Semantics: Semantic Structure Parsing for Table-Centric Document Analysis
- 构建表为中心的语义解析框架,挖掘表格与上下文关联。
- 在近4000页真实文档上达到90%以上精确率和F1值。
- 适合需要精准提取表格语义信息的研究者与工程师。
文档是信息与知识的核心载体,在金融、医疗和科研等领域广泛应用。表格作为结构化数据的主要媒介,包含关键信息,是文档中最为重要的组成部分。现有研究多聚焦于布局分析、表格检测和数据抽取等表面任务,缺乏对表格及其上下文关系的深度语义解析,制约了跨段落数据理解与一致性分析等高级应用。为此,我们提出DOTABLER——一种面向表格的语义文档解析框架,旨在揭示表格与其上下文之间的深层语义联系。该框架基于自建数据集并采用领域特定预训练模型微调,整合完整解析流程,识别与表格语义相关的上下文片段。在此基础上,实现两个核心功能:表为中心的文档结构解析与领域特定表格检索,提供全面的表锚定语义分析与精准的语义相关表格提取。在近4000页、超过1000张表格的真实PDF文档上评估,DOTABLER在表-上下文语义分析与深层文档解析上优于GPT-4o等先进模型,精确率与F1分数均超90%。
原文摘要 · Abstract (English)
Documents are core carriers of information and knowl-edge, with broad applications in finance, healthcare, and scientific research. Tables, as the main medium for structured data, encapsulate key information and are among the most critical document components. Existing studies largely focus on surface-level tasks such as layout analysis, table detection, and data extraction, lacking deep semantic parsing of tables and their contextual associations. This limits advanced tasks like cross-paragraph data interpretation and context-consistent analysis. To address this, we propose DOTABLER, a table-centric semantic document parsing framework designed to uncover deep semantic links between tables and their context. DOTABLER leverages a custom dataset and domain-specific fine-tuning of pre-trained models, integrating a complete parsing pipeline to identify context segments semantically tied to tables. Built on this semantic understanding, DOTABLER implements two core functionalities: table-centric document structure parsing and domain-specific table retrieval, delivering comprehensive table-anchored semantic analysis and precise extraction of semantically relevant tables. Evaluated on nearly 4,000 pages with over 1,000 tables from real-world PDFs, DOTABLER achieves over 90% Precision and F1 scores, demonstrating superior performance in table-context semantic analysis and deep document parsing compared to advanced models such as GPT-4o.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。