专为植物抗逆文献设计的开源语言模型,精准提取生物关系。
PlantDeBERTa: An Open Source Language Model for Plant Science
- 基于DeBERTa架构,结合规则与本体增强后处理
- 在鹰嘴豆抗逆文献上实现高精度实体识别与关系抽取
- 适合植物基因组学、农学知识发现研究者使用
Transformer语言模型在生物医学领域取得突破,但植物科学仍缺乏针对性工具。本文提出PlantDeBERTa,一个基于DeBERTa架构、针对鹰嘴豆(Lens culinaris)对多种非生物和生物胁迫响应文献的高性能开源语言模型。该模型在专家标注的摘要语料库上微调,采用分层本体标注框架(作物本体),覆盖分子、生理、生化和农艺维度。结合基于规则的后处理与本体引导的实体归一化,可精准捕捉生物学意义明确的关系。模型在多种实体类型上表现出强泛化能力,验证了低资源科学领域中稳健领域适配的可行性。我们公开发布模型,推动计算植物科学的透明协作与跨学科创新。
原文摘要 · Abstract (English)
The rapid advancement of transformer-based language models has catalyzed breakthroughs in biomedical and clinical natural language processing; however, plant science remains markedly underserved by such domain-adapted tools. In this work, we present PlantDeBERTa, a high-performance, open-source language model specifically tailored for extracting structured knowledge from plant stress-response literature. Built upon the DeBERTa architecture-known for its disentangled attention and robust contextual encoding-PlantDeBERTa is fine-tuned on a meticulously curated corpus of expert-annotated abstracts, with a primary focus on lentil (Lens culinaris) responses to diverse abiotic and biotic stressors. Our methodology combines transformer-based modeling with rule-enhanced linguistic post-processing and ontology-grounded entity normalization, enabling PlantDeBERTa to capture biologically meaningful relationships with precision and semantic fidelity. The underlying corpus is annotated using a hierarchical schema aligned with the Crop Ontology, encompassing molecular, physiological, biochemical, and agronomic dimensions of plant adaptation. PlantDeBERTa exhibits strong generalization capabilities across entity types and demonstrates the feasibility of robust domain adaptation in low-resource scientific fields.By providing a scalable and reproducible framework for high-resolution entity recognition, PlantDeBERTa bridges a critical gap in agricultural NLP and paves the way for intelligent, data-driven systems in plant genomics, phenomics, and agronomic knowledge discovery. Our model is publicly released to promote transparency and accelerate cross-disciplinary innovation in computational plant science.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。