构建医学文本嵌入模型与评估体系,提升医疗应用效果。
Towards Domain Specification of Embedding Models in Medicine
- 基于多源自监督对比学习,微调GTE模型以覆盖多样医学语料。
- 在51项任务中表现超越现有模型,尤其在术语与语义多样性场景下优势明显。
- 适合医疗AI研究者、临床决策系统开发者使用。
医学文本嵌入模型是临床决策支持、生物信息检索和医疗问答等应用的基础,但存在两大短板:一是多数模型仅在有限医学与生物数据上训练,且方法陈旧,难以捕捉实际中术语与语义的多样性;二是现有评估不足,主流基准无法泛化至真实医疗任务全谱。为此,我们采用在多源医学语料上通过自监督对比学习广泛微调的GTE模型(MEDTE),生成稳健的医学文本嵌入。同时,提出涵盖51项任务的综合性评估套件,覆盖分类、聚类、成对分类与检索,参照大规模文本嵌入基准(MTEB)但针对医学文本特点定制。结果表明,该方法不仅建立可靠评估框架,所产嵌入在各项任务中均持续优于现有最优模型。
原文摘要 · Abstract (English)
Medical text embedding models are foundational to a wide array of healthcare applications, ranging from clinical decision support and biomedical information retrieval to medical question answering, yet they remain hampered by two critical shortcomings. First, most models are trained on a narrow slice of medical and biological data, beside not being up to date in terms of methodology, making them ill suited to capture the diversity of terminology and semantics encountered in practice. Second, existing evaluations are often inadequate: even widely used benchmarks fail to generalize across the full spectrum of real world medical tasks. To address these gaps, we leverage MEDTE, a GTE model extensively fine-tuned on diverse medical corpora through self-supervised contrastive learning across multiple data sources, to deliver robust medical text embeddings. Alongside this model, we propose a comprehensive benchmark suite of 51 tasks spanning classification, clustering, pair classification, and retrieval modeled on the Massive Text Embedding Benchmark (MTEB) but tailored to the nuances of medical text. Our results demonstrate that this combined approach not only establishes a robust evaluation framework but also yields embeddings that consistently outperform state of the art alternatives in different tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。