arXiv:2604.16241cs.CLcs.AI2026-04

评测大模型对动物知识的掌握程度,发现其在生物多样性领域存在系统性短板。

BAGEL: Benchmarking Animal Knowledge Expertise in Language Models

论文配图:BAGEL: Benchmarking Animal Knowledge Expertise in Language Models
图 1 · 摘自论文原文
  • 基于科学文献与数据库构建闭卷问答集,统一评估模型动物知识
  • 覆盖分类、形态、分布等7类动物知识,支持细粒度分析
  • 适合研究模型在生物多样性任务中的泛化能力与可靠性

大型语言模型在通用知识和推理基准上表现优异,但其在统一闭卷评估协议下对专业动物知识的掌握程度仍不明确。我们提出BAGEL,一个用于评估语言模型动物知识专长的基准。BAGEL源自多样化的科学与参考资源,包括bioRxiv、Global Biotic Interactions、Xeno-canto和Wikipedia,结合人工筛选示例与自动生成的闭卷问答对构建。该基准涵盖动物知识的多个方面:分类学、形态学、栖息地、行为、发声、地理分布及物种互作。通过聚焦闭卷评估,BAGEL在推理时无需外部检索即可测量模型的动物相关知识。此外,该基准支持跨来源领域、分类群和知识类别进行细粒度分析,可更精确刻画模型优势与系统性失败模式。BAGEL为研究语言模型在特定领域知识泛化问题提供了新测试平台,并有助于提升其在生物多样性相关应用中的可靠性。

原文摘要 · Abstract (English)

Large language models have shown strong performance on broad-domain knowledge and reasoning benchmarks, but it remains unclear how well language models handle specialized animal-related knowledge under a unified closed-book evaluation protocol. We introduce BAGEL, a benchmark for evaluating animal knowledge expertise in language models. BAGEL is constructed from diverse scientific and reference sources, including bioRxiv, Global Biotic Interactions, Xeno-canto, and Wikipedia, using a combination of curated examples and automatically generated closed-book question-answer pairs. The benchmark covers multiple aspects of animal knowledge, including taxonomy, morphology, habitat, behavior, vocalization, geographic distribution, and species interactions. By focusing on closed-book evaluation, BAGEL measures animal-related knowledge of models without external retrieval at inference time. BAGEL further supports fine-grained analysis across source domains, taxonomic groups, and knowledge categories, enabling a more precise characterization of model strengths and systematic failure modes. Our benchmark provides a new testbed for studying domain-specific knowledge generalization in language models and for improving their reliability in biodiversity-related applications.

动物知识闭卷评测生物多样性大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。