首个融合空间与语义信息的文档建模格式,支持多页文档统一表示。
S2Doc -- Spatial-Semantic Document Format
- 提出融合空间布局与语义结构的统一文档数据格式
- 支持多页文档及多种任务建模,兼容主流文档处理方法
- 解决现有格式不兼容问题,适合文档理解与结构化任务
文档是存储和共享信息的常见方式,表格在其中占据重要地位。然而,目前缺乏对文档(尤其是表格)建模的统一标准,导致各类研究采用不同的数据结构和格式,彼此不兼容。现有模型通常只关注文档的空间或语义结构,忽视另一维度。为此,我们提出 S2Doc,一种灵活的文档与表格建模数据结构,将空间与语义信息整合于单一格式中。该格式易于扩展至新任务,支持多页文档,并兼容大多数文档与表格建模方法。据我们所知,S2Doc 是首个同时实现这些特性的统一格式。
原文摘要 · Abstract (English)
Documents are a common way to store and share information, with tables being an important part of many documents. However, there is no real common understanding of how to model documents and tables in particular. Because of this lack of standardization, most scientific approaches have their own way of modeling documents and tables, leading to a variety of different data structures and formats that are not directly compatible. Furthermore, most data models focus on either the spatial or the semantic structure of a document, neglecting the other aspect. To address this, we developed S2Doc, a flexible data structure for modeling documents and tables that combines both spatial and semantic information in a single format. It is designed to be easily extendable to new tasks and supports most modeling approaches for documents and tables, including multi-page documents. To the best of our knowledge, it is the first approach of its kind to combine all these aspects in a single format.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。