arXiv:2510.02830cs.CLcs.AI2025-10被引 2

LLM在物种分类准但评估风险差,需人类把关。

Evaluating Large Language Models for IUCN Red List Species Information

  • 用5个主流模型测试2万1千955个物种的分类与评估能力。
  • 分类准确率达94.9%,但保护状态判断仅27.2%正确。
  • 模型偏好大型动物,可能加剧保护不公,适合辅助而非替代专家。

大型语言模型(LLMs)正被快速应用于保护领域以应对生物多样性危机,但其在物种评估中的可靠性尚不明确。本研究系统验证了五个领先模型在21,955个物种上的表现,涵盖IUCN红色名录四个核心评估维度:分类学、保护状况、分布和威胁。关键发现是:模型在分类任务中表现优异(准确率94.9%),但在保护推理方面严重不足(保护状态评估准确率仅27.2%)。这一知识-推理差距在所有模型中均存在,表明其源于架构限制,而不仅是数据问题。此外,模型对标志性脊椎动物存在系统性偏倚,可能加剧现有保护不平等。研究明确指出:LLM适用于信息检索,但涉及判断的决策必须由人类主导。建议采用混合模式,让模型增强专家能力,而风险评估与政策制定仍由人类专家全权负责。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are rapidly being adopted in conservation to address the biodiversity crisis, yet their reliability for species evaluation is uncertain. This study systematically validates five leading models on 21,955 species across four core IUCN Red List assessment components: taxonomy, conservation status, distribution, and threats. A critical paradox was revealed: models excelled at taxonomic classification (94.9%) but consistently failed at conservation reasoning (27.2% for status assessment). This knowledge-reasoning gap, evident across all models, suggests inherent architectural constraints, not just data limitations. Furthermore, models exhibited systematic biases favoring charismatic vertebrates, potentially amplifying existing conservation inequities. These findings delineate clear boundaries for responsible LLM deployment: they are powerful tools for information retrieval but require human oversight for judgment-based decisions. A hybrid approach is recommended, where LLMs augment expert capacity while human experts retain sole authority over risk assessment and policy.

大模型保护评估IUCN偏见检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。