用知识图谱提升生物注释优先级,减少专家人工审核负担。
Plausibility-Driven Prioritization of Candidate Biomedical Annotations

- 基于生物知识图谱训练关系特异的分类器,提升注释可信度。
- 新方法在5个大型图谱上平均提升平衡准确率5.8%。
- 适合需要高效筛选生物注释的科研人员和数据库构建者。
生物医学知识的快速积累使得自动注释的验证成为生物信息注释中的主要瓶颈。尽管计算方法能快速生成大量候选注释,但判断其生物学有效性仍需耗费高昂成本的专家评审。因此,在人工注释前对候选注释进行优先排序成为关键挑战。本文提出一种利用生物知识图谱(bioKGs)评估候选注释合理性并引导专家注释的框架。从知识图谱嵌入出发,采用基于社区的负采样策略训练关系特异的二分类器,获得可靠的置信度估计。进一步提出一组合理性度量,结合分类器置信度、可靠性以及同一生物实体对间其他关系提供的语义上下文。与传统置信度估计不同,该方法显式考虑同一实体对可能共存的多种生物学关系。在五个大型生物知识图谱上的实验表明,所提出的负采样策略显著提升了分类器鲁棒性,平均平衡准确率提高5.8%。此外,合理性度量优于仅依赖分类器置信度的方法,能更有效地为专家评审筛选候选注释。结果表明,利用生物知识图谱可提升人工智能辅助生物注释的效率,同时保持专家对最终注释的控制权。
原文摘要 · Abstract (English)
The rapid growth of biomedical knowledge has made the validation of automatically generated biological annotations a major bottleneck in biomedical curation. While computational methods can rapidly produce large numbers of candidate annotations, determining which are biologically valid still requires costly expert review. Prioritizing these candidates before manual curation has therefore become a fundamental challenge. Machine learning techniques can support this process by exploiting biomedical knowledge graphs (bioKGs), which capture biological entities and their functional associations. In this work, we propose a framework that leverages bioKGs to estimate the plausibility of candidate annotations and guide expert curation. Starting from knowledge graph embeddings, we train relation-specific binary classifiers using a community-based negative sampling strategy to obtain reliable confidence estimates. We then introduce a family of plausibility measures that combine classifier confidence, classifier reliability, and the semantic context provided by alternative relationships involving the same pair of biological entities. Unlike conventional confidence estimation, the proposed approach explicitly accounts for multiple biologically meaningful relations that may coexist between the same entities. Experimental results on five large bioKGs demonstrate that the proposed negative sampling strategy consistently improves classifier robustness, increasing balanced accuracy by an average of 5.8%. Moreover, the plausibility measures outperform classifier confidence alone, enabling more effective prioritization of candidate annotations for expert review. Overall, our results show that the use of bioKGs improves the efficiency of AI-assisted biomedical curation while preserving expert control over the final annotation assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。