用大模型分析错误,引导医学关系分类模型逐步提升
Error-Aware Curriculum Learning for Biomedical Relation Classification
- 让GPT-4o分析学生模型的错误类型,生成针对性修正建议
- 在5个数据集上取得4个新最好结果,最高提升达6.2%准确率
- 适合需要高精度医学知识提取的研究者和临床应用开发者
生物医学文本中的关系分类(RC)对构建知识图谱及药物重定位、临床决策等应用至关重要。本文提出一种误差感知的师生框架,利用GPT-4o对基线学生模型的预测失败进行分析,识别错误类型,分配难度分,并生成改写句子与基于知识图谱的增强建议。这些增强标注用于通过指令微调训练首个学生模型。该模型再对更广数据集进行标注,附带难度分与增强输入。随后在按难度排序的数据集上,通过课程学习训练第二个学生模型,实现稳健渐进式学习。我们还从PubMed摘要构建了一个异构生物医学知识图谱,支持上下文感知的RC。本方法在5个蛋白质互作(PPI)数据集中的4个以及药物相互作用(DDI)数据集上达到新最优性能,在ChemProt数据集上保持竞争力。
原文摘要 · Abstract (English)
Relation Classification (RC) in biomedical texts is essential for constructing knowledge graphs and enabling applications such as drug repurposing and clinical decision-making. We propose an error-aware teacher--student framework that improves RC through structured guidance from a large language model (GPT-4o). Prediction failures from a baseline student model are analyzed by the teacher to classify error types, assign difficulty scores, and generate targeted remediations, including sentence rewrites and suggestions for KG-based enrichment. These enriched annotations are used to train a first student model via instruction tuning. This model then annotates a broader dataset with difficulty scores and remediation-enhanced inputs. A second student is subsequently trained via curriculum learning on this dataset, ordered by difficulty, to promote robust and progressive learning. We also construct a heterogeneous biomedical knowledge graph from PubMed abstracts to support context-aware RC. Our approach achieves new state-of-the-art performance on 4 of 5 PPI datasets and the DDI dataset, while remaining competitive on ChemProt.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。