arXiv:2505.13220cs.CL2025-05ACL被引 10

首个面向种子科学的多任务评测基准,助力大模型辅助育种

SeedBench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science

  • 构建跨任务种子科学评测基准,模拟现代育种流程
  • 评估26个主流大模型,揭示其与真实育种需求的显著差距
  • 适合农业AI研究者、育种专家及大模型应用开发者

种子科学对现代农业至关重要,直接影响作物产量和全球粮食安全。然而,跨学科复杂性高、研发成本大且回报有限,导致专业人才短缺和技术支持不足。尽管大语言模型在多个领域展现潜力,但在种子科学中的应用受限于数字资源匮乏、基因-性状关系复杂及缺乏标准化评测基准。为此,我们提出SeedBench——首个专为种子科学设计的多任务基准。该基准由领域专家共同开发,聚焦种子育种,模拟现代育种关键环节。我们对26个领先大模型(含商用、开源及领域微调模型)进行了全面评估。结果不仅揭示了大模型能力与实际育种问题之间的显著差距,也为大模型在种子设计研究中奠定了基础。

原文摘要 · Abstract (English)

Seed science is essential for modern agriculture, directly influencing crop yields and global food security. However, challenges such as interdisciplinary complexity and high costs with limited returns hinder progress, leading to a shortage of experts and insufficient technological support. While large language models (LLMs) have shown promise across various fields, their application in seed science remains limited due to the scarcity of digital resources, complex gene-trait relationships, and the lack of standardized benchmarks. To address this gap, we introduce SeedBench -- the first multi-task benchmark specifically designed for seed science. Developed in collaboration with domain experts, SeedBench focuses on seed breeding and simulates key aspects of modern breeding processes. We conduct a comprehensive evaluation of 26 leading LLMs, encompassing proprietary, open-source, and domain-specific fine-tuned models. Our findings not only highlight the substantial gaps between the power of LLMs and the real-world seed science problems, but also make a foundational step for research on LLMs for seed design.

种子科学大模型评测农业AI多任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。