arXiv:2409.10016cs.CLcs.AI2024-09被引 1

首个涵盖多种学术文本结构的解析数据集,助力论文智能处理。

AceParse: A Comprehensive Dataset with Diverse Structured Texts for Academic Literature Parsing

  • 构建首个覆盖公式、表格、列表等多样结构的学术文本数据集
  • 新模型在F1和相似度上分别提升4.1%和5%,超越现有水平
  • 适合研究文献解析、多模态模型与AI辅助写作的开发者使用

随着以数据为中心的AI发展,研究重点从模型驱动转向提升数据质量。学术文献作为重要数据源,主要以PDF形式存储,需先解析为可处理文本。然而,由于缺乏涵盖多种文本结构的数据集,学术文献中复杂结构的解析仍具挑战。本文提出AceParse,首个专为支持多样化结构文本解析设计的综合性数据集,涵盖公式、表格、列表、算法及含数学表达的句子。基于该数据集,我们微调了名为AceParser的多模态模型,能准确解析学术文献中的各类结构化内容。该模型在F1分数上较此前最优方法提升4.1%,在Jaccard相似度上提升5%,展现出多模态模型在学术文献解析中的潜力。数据集已公开于https://github.com/JHW5981/AceParse。

原文摘要 · Abstract (English)

With the development of data-centric AI, the focus has shifted from model-driven approaches to improving data quality. Academic literature, as one of the crucial types, is predominantly stored in PDF formats and needs to be parsed into texts before further processing. However, parsing diverse structured texts in academic literature remains challenging due to the lack of datasets that cover various text structures. In this paper, we introduce AceParse, the first comprehensive dataset designed to support the parsing of a wide range of structured texts, including formulas, tables, lists, algorithms, and sentences with embedded mathematical expressions. Based on AceParse, we fine-tuned a multimodal model, named AceParser, which accurately parses various structured texts within academic literature. This model outperforms the previous state-of-the-art by 4.1% in terms of F1 score and by 5% in Jaccard Similarity, demonstrating the potential of multimodal models in academic literature parsing. Our dataset is available at https://github.com/JHW5981/AceParse.

文献解析多模态数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。