arXiv:2605.28521cs.CL2026-05被引 2

多语言医学文本检索模型,提升跨语言疾病实体识别效果。

ClinicalEncoder26AM: A Multlilingual Diagnosable ColBERT Model; Evidences from the MultiClinNER Shared Task

  • 基于临床语义空间构建多级对齐的可诊断模型
  • 在多语言医疗命名实体任务中达顶尖召回率与F1分数
  • 适合需要高效跨语言医疗信息抽取的研究与应用

ClinicalEncoder26AM 是一个用于临床和生物医学文本的多语言可诊断 ColBERT 模型,其词元级语义与 ClinicalMap25 对齐,后者是受 BioLORD-2023 启发并融合合成及标注监督的临床潜在空间。该模型在 BGE-M3 基础上进行后训练,结合合成临床病历、医患对话及 MedMentions 等标注资源,通过多适配器蒸馏方式同时优化命名实体级别与句子级别表征,并采用 ColBERT 风格的检索目标。在 MultiClinNER 共享任务中,我们将其微调为 BIO 标记器,使用轻量级两层 CNN 头增强局部边界检测。系统保持简洁,大多数文档可在单个 8192 令牌窗口内处理,实现最先进的多语言实体召回率,在所有实体类型和语言上的字符加权 F1 得分位列前五。训练曲线显示,该模型比基础 M3 模型显著更数据高效,验证了其临床后训练对下游信息提取的有效性。模型可在 https://huggingface.co/Parallia/ClinicalEncoder26AM-Diagnosable-Colbert-L2-for-multilingual-medical-texts 下载。

原文摘要 · Abstract (English)

ClinicalEncoder26AM is a multilingual Diagnosable ColBERT for clinical and biomedical texts, which aligns at multiple levels its token-level semantic with ClinicalMap25, a clinical latent space inspired by BioLORD-2023 and enriched with synthetic and annotated supervision. The post-training recipe builds upon BGE-M3, and combines synthetic clinical notes, patient--doctor conversations, and annotated resources such as MedMentions, while considering both named-entity-level and sentence-level representations in a multi-adapter distillation, along with a ColBERT-style retrieval objective. In this system demonstration paper, we evaluate the model in the MultiClinNER shared task by finetuning it as a BIO tagger for patient symptoms, disorders, and procedure spans, using a lightweight two-layer CNN head to improve local boundary detection. The resulting system remains simple, processes most documents in a single 8192-token window, and achieves state-of-the-art multilingual entity recall, while achieving Top 5 overall across all entity types and languages in Character-weighted F1 scores. Training curves further show that ClinicalEncoder26AM is markedly more data-efficient than the base M3 model, supporting the usefulness of its clinical post-training for downstream information extraction. The model can be downloaded on https://huggingface.co/Parallia/ClinicalEncoder26AM-Diagnosable-Colbert-L2-for-multilingual-medical-texts

多语言医学文本实体识别可诊断模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。