arXiv:2507.19315cs.CL2025-07中稿 · ISMB 2026被引 10

无需特定训练,自动识别生物表型概念

AutoPCR: Automated Phenotype Concept Recognition by Prompting

  • 用提示工程实现跨本体自动识别
  • 在多个数据集上表现最优且稳定
  • 适合需要快速适配新领域研究者

表型概念识别是生物医学文本挖掘的基础任务。现有方法或需针对本体进行特定训练,难以适应多样的文本风格和不断演化的术语;或依赖通用大模型,缺乏必要的领域知识。为此,我们提出AutoPCR,一种基于提示的表型概念识别方法,可自动泛化至新本体和未见数据,无需本体特定训练。为进一步提升性能,引入可选的自监督训练策略。实验表明,AutoPCR在多个数据集上取得最佳平均性能与最强鲁棒性。消融与迁移实验验证其归纳能力及对新本体的泛化能力。代码已开源:https://github.com/yctao7/AutoPCR。联系邮箱:[email protected]

原文摘要 · Abstract (English)

Motivation: Phenotype concept recognition (CR) is a fundamental task in biomedical text mining. However, existing methods either require ontology-specific training, making them struggle to generalize across diverse text styles and evolving biomedical terminology, or depend on general-purpose large language models (LLMs) that lack necessary domain knowledge. Results: To address these limitations, we propose AutoPCR, a prompt-based phenotype CR method designed to automatically generalize to new ontologies and unseen data without ontology-specific training. To further boost performance, we also introduce an optional self-supervised training strategy. Experiments show that AutoPCR achieves the best average and most robust performance across datasets. Further ablation and transfer studies demonstrate its inductive capability and generalizability to new ontologies. Availability and Implementation: Our code is available at https://github.com/yctao7/AutoPCR. Contact: [email protected]

表型识别提示工程零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。