用大模型自动从网页提取植物形态特征,效率远超人工。
Fully automatic extraction of morphological traits from the Web: utopia or reality?
- 用大语言模型解析网页文本,自动提取植物性状数据
- 对三组人工矩阵复现,超过一半数据对识别成功,F1超75%
- 适合生态学、植物分类研究者快速构建性状数据库
植物形态性状是理解物种在生态系统中作用的基础。但为中等数量物种整理性状信息需专家多年时间。与此同时,大量物种描述文本在线存在,但缺乏结构,难以规模化利用。为此,我们提出利用大语言模型(LLMs)处理非结构化文本中的植物性状信息,无需人工标注。通过自动复现三个手工构建的物种-性状矩阵,我们的方法在超过一半的物种-性状对中成功识别出数值,F1得分超过75%。结果表明,借助大模型的信息抽取能力,目前可实现从非结构化网络文本大规模构建结构化性状数据库,瓶颈在于覆盖目标性状的文本是否充足。
原文摘要 · Abstract (English)
Plant morphological traits, their observable characteristics, are fundamental to understand the role played by each species within their ecosystem. However, compiling trait information for even a moderate number of species is a demanding task that may take experts years to accomplish. At the same time, massive amounts of information about species descriptions is available online in the form of text, although the lack of structure makes this source of data impossible to use at scale. To overcome this, we propose to leverage recent advances in large language models (LLMs) and devise a mechanism for gathering and processing information on plant traits in the form of unstructured textual descriptions, without manual curation. We evaluate our approach by automatically replicating three manually created species-trait matrices. Our method managed to find values for over half of all species-trait pairs, with an F1-score of over 75%. Our results suggest that large-scale creation of structured trait databases from unstructured online text is currently feasible thanks to the information extraction capabilities of LLMs, being limited by the availability of textual descriptions covering all the traits of interest.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。