arXiv:2511.04573cs.LG2025-11被引 1

用大模型自动提取文献中的物种分布数据,大幅加速生态研究

ARETE: an R package for Automated REtrieval from TExt with large language models

  • 基于大语言模型构建R包,自动化从文本中提取物种出现记录
  • 对100种蜘蛛的分析显示,新数据使已知分布范围扩大了1000倍
  • 适合需要海量物种数据的保护规划与风险评估研究者使用

当前生态保护面临关键物种数据缺失的瓶颈,尤其缺乏可机器读取的出现数据。受人类活动影响,科研人员需快速获取和处理大量新信息。科学论文与灰色文献虽含关键数据,但多为非结构化文本,依赖人工提取效率低下。本文提出ARETE R包,一个开源工具,利用大语言模型(如chatGPT API)实现物种出现数据的自动化提取与验证,涵盖光学字符识别、异常值检测到表格输出的全流程。通过与人工标注对比验证其准确性。以100种蜘蛛为例,自动提取数据使平均已知分布范围扩展了三个数量级,揭示了过去未被记录的分布区域,对空间保护规划与灭绝风险评估具有重要意义。ARETE显著提升了以往难以获取的数据访问速度,为资源优先分配和项目规划提供了支持。

原文摘要 · Abstract (English)

1. A hard stop for the implementation of rigorous conservation initiatives is our lack of key species data, especially occurrence data. Furthermore, researchers have to contend with an accelerated speed at which new information must be collected and processed due to anthropogenic activity. Publications ranging from scientific papers to gray literature contain this crucial information but their data are often not machine-readable, requiring extensive human work to be retrieved. 2. We present the ARETE R package, an open-source software aiming to automate data extraction of species occurrences powered by large language models, namely using the chatGPT Application Programming Interface. This R package integrates all steps of the data extraction and validation process, from Optical Character Recognition to detection of outliers and output in tabular format. Furthermore, we validate ARETE through systematic comparison between what is modelled and the work of human annotators. 3. We demonstrate the usefulness of the approach by comparing range maps produced using GBIF data and with those automatically extracted for 100 species of spiders. Newly extracted data allowed to expand the known Extent of Occurrence by a mean three orders of magnitude, revealing new areas where the species were found in the past, which mayhave important implications for spatial conservation planning and extinction risk assessments. 4. ARETE allows faster access to hitherto untapped occurrence data, a potential game changer in projects requiring such data. Researchers will be able to better prioritize resources, manually verifying selected species while maintaining automated extraction for the majority. This workflow also allows predicting available bibliographic data during project planning.

生态数据大模型自动化提取物种分布

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。