arXiv:2501.18287cs.CLcs.AI2025-01中稿 · the NLP4Ecology Wo…被引 9

用大模型从入侵生物学论文中自动提取物种、地点、生境和生态系统信息。

Mining for Species, Locations, Habitats, and Ecosystems from Scientific Papers in Invasion Biology: A Large-Scale Exploratory Study with Large Language Models

  • 直接使用通用大模型,无需领域微调即可识别生态实体。
  • 在未标注数据上实现对物种与地点的高精度抽取,验证了大模型潜力。
  • 为生态学家提供自动化知识提取工具,助力入侵管理与保护决策。

本文开展一项探索性研究,利用大语言模型(LLMs)从入侵生物学文献中挖掘关键生态实体。重点包括物种名称、地理位置、相关生境及生态系统信息,这些信息对理解物种扩散、预测未来入侵以及支持保护工作至关重要。传统文本挖掘方法常受限于生态术语的复杂性和文本中微妙的语言模式。本研究采用未经领域微调的通用大模型,在未标注数据上实现了对物种和地点的高效抽取,揭示了大模型在生态信息提取中的潜力与局限。该工作为构建更先进的自动化知识提取工具奠定了基础,可助力研究人员和实践者更好地理解和管理生物入侵问题。

原文摘要 · Abstract (English)

This paper presents an exploratory study that harnesses the capabilities of large language models (LLMs) to mine key ecological entities from invasion biology literature. Specifically, we focus on extracting species names, their locations, associated habitats, and ecosystems, information that is critical for understanding species spread, predicting future invasions, and informing conservation efforts. Traditional text mining approaches often struggle with the complexity of ecological terminology and the subtle linguistic patterns found in these texts. By applying general-purpose LLMs without domain-specific fine-tuning, we uncover both the promise and limitations of using these models for ecological entity extraction. In doing so, this study lays the groundwork for more advanced, automated knowledge extraction tools that can aid researchers and practitioners in understanding and managing biological invasions.

大模型信息抽取生态学入侵物种

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。