arXiv:2603.07050cs.IR2026-03

用大模型自动构建科学数据库,效率比人工高90%以上

Leveraging Large Language Models for Automated Scalable Development of Open Scientific Databases

  • 结合关键词检索、API获取与大模型分类,实现自动化数据采集
  • 在农业领域测试中,与专家数据库重合率达90%
  • 框架可跨领域扩展,适合需要大规模数据建设的研究者

随着在线科学文献的指数增长,识别可靠领域的特定数据变得愈发重要但也极为困难。手动收集和筛选特定领域科学文献不仅耗时费力,还易出错且不一致。为促进自动化数据收集,本文提出一个基于大语言模型(LLMs)的网络工具,用于自动化、可扩展地构建开放科学数据库。该工具采用统一的自动化框架,融合关键词查询、支持API的数据获取及大模型驱动的文本分类,从多个可信数据源和搜索引擎并行获取数据,形成整合数据集。随后,针对每个关键词查询定制提示词,调用大模型对数据进行筛选,提取与科学问题相关的条目。该方法在农业与作物产量相关的一系列领域任务中进行了测试,结果表明其与小型专家标注数据库的重合度达90%,证明该工具能显著减少人工工作量。此外,该框架具有可扩展性和领域无关性,适用于不同领域构建可扩展的开放科学数据库。

原文摘要 · Abstract (English)

With the exponential increase in online scientific literature, identifying reliable domain-specific data has become increasingly important but also very challenging. Manual data collection and filtering for domain-specific scientific literature is not only time-consuming but also labor-intensive and prone to errors and inconsistencies. To facilitate automated data collection, the paper introduces a web-based tool that leverages Large Language Models (LLMs) for automated and scalable development of open scientific databases. More specifically, the tool is based on an automated and unified framework that combines keyword-based querying, API-enabled data retrieval, and LLM-powered text classification to construct domain-specific scientific databases. Data is collected from multiple reliable data sources and search engines using a parallel querying technique to construct a combined unified dataset. The dataset is subsequently filtered using LLMs queried with prompts tailored for each keyword-based query to extract the relevant data to a scientific query of interest. The approach was tested across a set of variable keyword-based searches for different domain-specific tasks related to agriculture and crop yield. The results and analysis show 90\% overlap with small domain expert-curated databases, suggesting that the proposed tool can be used to significantly reduce manual workload. Furthermore, the proposed framework is both scalable and domain-agnostic and can be applied across diverse fields for building scalable open scientific databases.

大模型应用科学数据库自动化采集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。