用大模型生成医学实体链接训练数据,少用人标注也能超好用
SynCABEL: Synthetic Contextualized Augmentation for Biomedical Entity Linking
- 用大模型自动造带上下文的医学实体训练样本
- 在三个语种数据集上达最新水平,用60%更少标注数据就达标
- 新评测法发现它更准识别临床有效结果,适合医疗AI研究者
我们提出SynCABEL(合成上下文增强的生物医学实体链接框架),解决监督式生物医学实体链接(BEL)的核心瓶颈——专家标注数据稀缺问题。SynCABEL利用大语言模型为目标知识库中所有候选概念生成丰富上下文的合成训练样本,实现无需人工标注的大范围监督。实验表明,SynCABEL结合解码器仅模型与引导推理,在三个常用多语言基准上取得新最优性能:MedMentions(英语)、QUAERO(法语)和SPACCC(西班牙语)。在数据效率评估中,SynCABEL仅需全量人工标注数据的40%即可达到相同性能,显著降低对耗时费力的专家标注依赖。此外,针对标准评估常因术语匹配忽略本体冗余而低估临床有效预测的问题,我们引入大模型作为评判者的新协议,分析显示SynCABEL大幅提升临床有效预测率。本文开源合成数据集、模型与代码,支持可复现与后续研究。
原文摘要 · Abstract (English)
We present SynCABEL (Synthetic Contextualized Augmentation for Biomedical Entity Linking), a framework that addresses a central bottleneck in supervised biomedical entity linking (BEL): the scarcity of expert-annotated training data. SynCABEL leverages large language models to generate context-rich synthetic training examples for all candidate concepts in a target knowledge base, providing broad supervision without manual annotation. We demonstrate that SynCABEL, when combined with decoder-only models and guided inference, establishes new state-of-the-art results across three widely used multilingual benchmarks: MedMentions for English, QUAERO for French, and SPACCC for Spanish. Evaluating data efficiency, we show that SynCABEL reaches the performance of full human supervision using up to 60% less annotated data, substantially reducing reliance on labor-intensive and costly expert labeling. Finally, acknowledging that standard evaluation based on exact code matching often underestimates clinically valid predictions due to ontology redundancy, we introduce an LLM-as-a-judge protocol. This analysis reveals that SynCABEL significantly improves the rate of clinically valid predictions. Our synthetic datasets, models, and code are released to support reproducibility and future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。