用大模型自动提取生态数据元信息,统一格式方便跨源检索。
Flexible metadata harvesting for ecology using large language models
- 用大模型从任意数据页提取结构化与非结构化元信息。
- 通过嵌入相似性与格式统一,识别数据集间关联关系。
- 适合需要整合多源生态数据的研究者快速构建知识图谱。
大型开放数据集可加速生态学研究,尤其通过复用多源数据获得新洞见。然而,研究人员需在不同数据平台间寻找合适数据集时,面临元信息缺失或标准不一的障碍。为此,我们开发了一种基于大语言模型(LLM)的元信息采集工具,能灵活提取任意数据集落地页中的元信息,并通过现有标准转换为用户自定义的统一格式。验证表明,该工具在提取结构化与非结构化元信息方面具有同等准确性,得益于其大模型后处理协议。此外,我们利用大模型计算嵌入相似性,并统一提取的元信息格式,实现规则化处理以识别数据集间的关联。该工具可灵活链接不同数据集的元信息,可用于本体构建或基于图的查询,例如在虚拟研究环境中发现相关生态与环境数据集。
原文摘要 · Abstract (English)
Large, open datasets can accelerate ecological research, particularly by enabling researchers to develop new insights by reusing datasets from multiple sources. However, to find the most suitable datasets to combine and integrate, researchers must navigate diverse ecological and environmental data provider platforms with varying metadata availability and standards. To overcome this obstacle, we have developed a large language model (LLM)-based metadata harvester that flexibly extracts metadata from any dataset's landing page, and converts these to a user-defined, unified format using existing metadata standards. We validate that our tool is able to extract both structured and unstructured metadata with equal accuracy, aided by our LLM post-processing protocol. Furthermore, we utilise LLMs to identify links between datasets, both by calculating embedding similarity and by unifying the formats of extracted metadata to enable rule-based processing. Our tool, which flexibly links the metadata of different datasets, can therefore be used for ontology creation or graph-based queries, for example, to find relevant ecological and environmental datasets in a virtual research environment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。