用大模型自动标注物种命名,准确率超人工,但生态文化类仍需改进。
Evaluation of the Automated Labeling Method for Taxonomic Nomenclature Through Prompt-Optimized Large Language Model
- 通过提示工程优化大模型,实现物种名分类自动化。
- 形态、地理类准确率达90%以上,生态行为类较低。
- 适合生物分类、数据库构建等需要批量标注的场景。
生物物种的学名由属名和种加词构成,后者常反映形态、生态、分布及文化背景。传统上,研究人员需手动审阅分类描述进行标注,处理大规模数据时耗时费力。本研究评估了利用大语言模型(LLM)结合提示工程进行自动物种名标注的可行性。基于Mammola等人构建的蜘蛛名称数据集,将优化后的LLM标注结果与人工标注对比。结果显示,模型在形态、地理类别中达到高准确率;但在生态与行为、现代与历史文化类别中表现较差,反映出对动物行为及文化语境理解的局限。未来工作将聚焦于通过少样本学习和检索增强生成技术提升性能,并拓展至更多生物类群的应用。
原文摘要 · Abstract (English)
Scientific names of organisms consist of a genus name and a species epithet, with the latter often reflecting aspects such as morphology, ecology, distribution, and cultural background. Traditionally, researchers have manually labeled species names by carefully examining taxonomic descriptions, a process that demands substantial time and effort when dealing with large datasets. This study evaluates the feasibility of automatic species name labeling using large language model (LLM) by leveraging their text classification and semantic extraction capabilities. Using the spider name dataset compiled by Mammola et al., we compared LLM-based labeling results-enhanced through prompt engineering-with human annotations. The results indicate that LLM-based classification achieved high accuracy in Morphology, Geography, and People categories. However, classification accuracy was lower in Ecology & Behavior and Modern & Past Culture, revealing challenges in interpreting animal behavior and cultural contexts. Future research will focus on improving accuracy through optimized few-shot learning and retrieval-augmented generation techniques, while also expanding the applicability of LLM-based labeling to diverse biological taxa.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。