自动识别文献中的生信软件和数据库名称,提升知识提取效率。
Automatic bioinformatic software named entity recognition from literature
- 融合词法与语义建模,利用SciBERT和掩码策略增强命名实体识别。
- 在两个基准数据集上显著优于bioNerDS2和ChatGPT等模型。
- 适合生物信息学研究者追踪工具使用趋势与领域发展动态。
生信软件和数据库是现代生命科学研究的核心,但其在科学文献中的提及常不一致且难以大规模系统识别。缺乏全面更新的生信资源目录阻碍了自动化生物医学知识抽取与数据分析的进展。本文提出SNAIL,一种混合命名实体识别框架,可从生物医学文本中自动识别生信软件与数据库(SW/DB)名称。SNAIL结合词法与语义建模:词法部分捕捉名称的拼写模式与上下文线索;语义部分利用SciBERT等基于Transformer的语言模型生成上下文嵌入,并采用显式标记掩码策略强化实体表示。通过整合引文提示提取与大语言模型辅助蒸馏的混合流程,构建了大规模训练语料。在两个独立基准数据集及真实科研文章上的评估显示,SNAIL显著优于现有方法,包括bioNerDS2等领域专用模型以及ChatGPT、Gemini、Grok和Claude等通用大模型。将SNAIL应用于大规模文献分析进一步揭示了不同期刊在生信子领域间的工具偏好差异。结果表明,SNAIL为科学文本中生信资源的精准、可扩展识别提供了有效方案,支持对工具使用与研究趋势的系统性元分析。
原文摘要 · Abstract (English)
Bioinformatics software and databases are essential components of modern life science research, yet their mentions in the scientific literature are often inconsistent and difficult to systematically identify at scale. The lack of a comprehensive and up-to-date catalog of bioinformatics resources hinders efforts toward automated biomedical knowledge extraction and streamlined data analysis. Here we present SNAIL, a hybrid named entity recognition framework designed to automatically identify bioinformatics software and database (SW/DB) names from biomedical texts. SNAIL integrates complementary lexical and semantic modeling strategies. The lexical component captures orthographic patterns and contextual cues characteristic of SW/DB names, while the semantic component leverages contextual embeddings generated by transformer-based language models such as SciBERT, combined with an explicit token-masking strategy to enhance entity-focused representations. A large training corpus was constructed automatically through a hybrid pipeline that integrates citation-hinted extraction with large language model-assisted distillation. Evaluation on two independent benchmark datasets and real-world research articles demonstrates that SNAIL substantially outperforms existing approaches, including domain-specific methods such as bioNerDS2 and general-purpose large language models such as ChatGPT, Gemini, Grok and Claude. Applying SNAIL to large-scale literature analysis further reveals distinct journal-level preferences across bioinformatics subfields. These results demonstrate that SNAIL provides an accurate and scalable solution for identifying bioinformatics resources in scientific texts and enables systematic meta-analysis of tool usage and research trends.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。