arXiv:2605.10032cs.CL2026-05

构建多物种植物标志基因证据推理基准,提升文献支撑的生物信息提取可信度。

PlantMarkerBench: A Multi-Species Benchmark for Evidence-Grounded Plant Marker Reasoning

论文配图:PlantMarkerBench: A Multi-Species Benchmark for Evidence-Grounded Plant Marker Reasoning
图 1 · 摘自论文原文
  • 基于文献全文构建跨物种证据标注体系,融合检索与人工校验
  • 覆盖4种植物5550条证据,验证模型对功能/间接等弱支持证据识别能力
  • 揭示主流模型在复杂生物学语境下易误判,适合生物信息与AI交叉研究者

细胞特异性标志基因是植物生物学的基础,但现有资源主要依赖手工整理数据库或高通量研究,未显式建模科学文献中的支持证据。我们提出PlantMarkerBench,一个用于评估从全篇生物论文中进行文献支撑的植物标志基因证据解释能力的多物种基准。该基准采用模块化标注流程,整合大规模文献检索、混合搜索、物种感知生物定位、结构化证据抽取及定向人工审查,覆盖拟南芥、玉米、水稻和番茄四种植物,共包含5,550条句子级证据实例,标注了标记-证据有效性、证据类型和支撑强度。定义两项任务:判断候选句是否为基因-细胞类型对提供有效标记证据,以及将证据分类为表达、定位、功能、间接或否定类别。我们在多种开源与闭源语言模型上测试不同提示策略,发现尽管前沿模型在直接表达证据上表现良好,但在功能、间接和弱支持证据上性能显著下降,证据类型混淆成为主要失败模式;开源模型在模糊生物学情境下还表现出更高的假阳性率。PlantMarkerBench为文献支撑的生物证据归因提供了具有挑战性且可复现的评估框架,支持未来可信科学信息提取与人工智能辅助植物生物学研究。

原文摘要 · Abstract (English)

Cell-type-specific marker genes are fundamental to plant biology, yet existing resources primarily rely on curated databases or high-throughput studies without explicitly modeling the supporting evidence found in scientific literature. We introduce PlantMarkerBench, a multi-species benchmark for evaluating literature-grounded plant marker evidence interpretation from full-text biological papers. PlantMarkerBench is constructed using a modular curation pipeline integrating large-scale literature retrieval, hybrid search, species-aware biological grounding, structured evidence extraction, and targeted human review. The benchmark spans four plant species -- Arabidopsis, maize, rice, and tomato -- and contains 5,550 sentence-level evidence instances annotated for marker-evidence validity, evidence type, and support strength. We define two benchmark tasks: determining whether a candidate sentence provides valid marker evidence for a gene-cell-type pair, and classifying the evidence into expression, localization, function, indirect, or negative categories. We benchmark diverse open-weight and closed-source language models across species and prompting strategies. Although frontier models achieve relatively strong performance on direct expression evidence, performance drops substantially on functional, indirect, and weak-support evidence, with evidence-type confusion emerging as a dominant failure mode. Open-weight models additionally exhibit elevated false-positive rates under ambiguous biological contexts. PlantMarkerBench provides a challenging and reproducible evaluation framework for literature-grounded biological evidence attribution and supports future research on trustworthy scientific information extraction and AI-assisted plant biology.

生物信息文献挖掘植物基因证据推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。