arXiv:2503.20179cs.CLcs.IR2025-03

用原型网络+低秩微调,高效识别癌症免疫治疗研究

ProtoBERT-LoRA: Parameter-Efficient Prototypical Finetuning for Immunotherapy Study Identification

  • 用原型嵌入约束类别可分性,结合LoRA实现小样本高效微调
  • 在71正例765负例测试集上达F1 0.624,召回率88.7%
  • 适合资源有限的生物医学文本分类任务,节省82%人工审阅

在基因表达综合数据库(GEO)中识别免疫检查点抑制剂(ICI)研究对癌症研究至关重要,但受限于语义模糊、极端类别不平衡和低资源环境下标注数据稀缺。本文提出ProtoBERT-LoRA,融合PubMedBERT、原型网络与低秩适应(LoRA)的混合框架,通过情景式原型训练实现类别可分嵌入,同时保留生物医学领域知识。数据集划分为:训练集(20正,20负),原型集(10正,10负),验证集(20正,200负),测试集(71正,765负)。在测试集上,ProtoBERT-LoRA获得F1分数0.624(精确率0.481,召回率0.887),优于规则系统、机器学习基线及微调的PubMedBERT。应用于44,287条未标注研究后,人工审查工作量减少82%。消融实验表明,原型与LoRA结合使性能较单独使用LoRA提升29%。

原文摘要 · Abstract (English)

Identifying immune checkpoint inhibitor (ICI) studies in genomic repositories like Gene Expression Omnibus (GEO) is vital for cancer research yet remains challenging due to semantic ambiguity, extreme class imbalance, and limited labeled data in low-resource settings. We present ProtoBERT-LoRA, a hybrid framework that combines PubMedBERT with prototypical networks and Low-Rank Adaptation (LoRA) for efficient fine-tuning. The model enforces class-separable embeddings via episodic prototype training while preserving biomedical domain knowledge. Our dataset was divided as: Training (20 positive, 20 negative), Prototype Set (10 positive, 10 negative), Validation (20 positive, 200 negative), and Test (71 positive, 765 negative). Evaluated on test dataset, ProtoBERT-LoRA achieved F1-score of 0.624 (precision: 0.481, recall: 0.887), outperforming the rule-based system, machine learning baselines and finetuned PubMedBERT. Application to 44,287 unlabeled studies reduced manual review efforts by 82%. Ablation studies confirmed that combining prototypes with LoRA improved performance by 29% over stand-alone LoRA.

生物医学小样本学习原型网络低秩微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。