arXiv:2502.04022cs.CL2025-02综述被引 2

用大模型和最佳最差法,从历史调查文本自动估算物种数量。

Quantification of Biodiversity from Historical Survey Text with LLM-based Best-Worst Scaling

  • 用大模型结合最佳最差法,将物种频率估计转为回归任务。
  • GPT-4和DeepSeek-V3与人类判断一致性高,结果可靠。
  • 比细粒度分类更省钱、更易扩展,适合大规模物种统计。

本研究评估了从历史调查文本中通过数量估计确定物种频次的方法。我们构建分类任务,最终发现该问题可通过大型语言模型(LLMs)结合最佳最差法(BWS)以回归任务形式有效建模。测试了Ministral-8B、DeepSeek-V3和GPT-4,结果显示后两者与人类判断及彼此间具有合理一致性。结论表明,该方法相较细粒度多分类方案更具成本效益且同样稳健,支持跨物种的自动化数量估算。

原文摘要 · Abstract (English)

In this study, we evaluate methods to determine the frequency of species via quantity estimation from historical survey text. To that end, we formulate classification tasks and finally show that this problem can be adequately framed as a regression task using Best-Worst Scaling (BWS) with Large Language Models (LLMs). We test Ministral-8B, DeepSeek-V3, and GPT-4, finding that the latter two have reasonable agreement with humans and each other. We conclude that this approach is more cost-effective and similarly robust compared to a fine-grained multi-class approach, allowing automated quantity estimation across species.

物种数量大模型历史数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。