arXiv:2608.14055cs.CL2026-08

将地质学长篇文献转为结构化数据,提升知识可计算性

HERMES: a multi-agent framework for structured knowledge extraction from ultra-long documents in geoscience

  • 多智能体框架整合文本、表格、图表与元信息,统一提取
  • 从55卷古生物专著中提取3.2万种化石实体与45万属性,F1超0.9
  • 无需重新训练即可跨古地磁、地球化学等领域的知识提取

地质学权威科学知识仍大量存在于传统专著与历史文献中,其非结构化文本和复杂版式阻碍了计算访问。我们提出HERMES,一个可扩展的多智能体框架,用于从超长科学文献中提取结构化数据。通过协调型大语言模型,HERMES在统一文档级提取流程中整合领域约束、验证规则与证据溯源,涵盖解析后的文本、表格、图表及其标题。应用于55卷《无脊椎古生物专著》,系统生成包含32,277个化石分类实体和451,878个属性的结构化数据库,并在线发布于https://treatise.geolex.org。各类群提取性能稳定(实体平均F1约0.90,属性约0.91),单卷效率较全人工基线提升约六倍。在古地磁学与地球化学领域的评估中,未进行额外模型训练即实现跨域迁移。本工作为将历史科学文献转化为面向FAIR原则的结构化数据提供了可行路径,为数据密集型学科与大规模知识融合构建可持续基础设施。

原文摘要 · Abstract (English)

Authoritative scientific knowledge in geoscience remains largely trapped in legacy monographs and historical literature, where unstructured text and complex layouts hinder computational access. We introduce HERMES, a scalable multi-agent framework that extracts structured data from ultra-long scientific documents. Using a coordinating large language model, HERMES integrates domain constraints, validation rules and evidence tracing within a unified document-level extraction process that incorporates parsed text, tables, figures and captions. Applied to the 55-volume Treatise on Invertebrate Paleontology, the system produced a structured database of 32,277 fossil taxonomic entities and 451,878 attributes, released online at https://treatise.geolex.org. Extraction performance remained stable across fossil groups (average F1 scores of approximately 0.90 for entities and 0.91 for attributes), improving per-volume efficiency approximately sixfold relative to the tested fully manual baseline. Evaluation in palaeomagnetism and geochemistry, conducted without additional model training, demonstrated transfer across distinct geoscience domains. This work provides a practical pathway to transform historical scientific literature into FAIR-oriented structured data, offering a sustainable infrastructure for data-intensive disciplines and large-scale knowledge integration.

知识提取地质学多智能体结构化数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。