arXiv:2608.29890cs.CL2026-09

首个英越双语医学实体识别语料库,支持跨语言医疗AI研究

En-ViMedNER: An English-Vietnamese Parallel Biomedical Corpus with UMLS Semantic Type Annotations

论文配图:En-ViMedNER: An English-Vietnamese Parallel Biomedical Corpus with UMLS Semantic Type Annotations
图 1 · 摘自论文原文
  • 通过自动翻译+专家校对构建双语医学实体语料
  • 越南语实体识别最佳模型F1达53.78,跨语言任务最高45.44
  • 专为越南语医学NLP研究设计,适配多语言医疗系统开发

生物医学命名实体识别(NER)是医疗AI应用的基础,如临床决策支持和医学信息提取。尽管已有包含统一医学语言系统(UMLS)标注的资源(如MedMentions),但目前尚无针对越南语的类似数据集。本文提出En-ViMedNER,首个英越双语平行生物医学NER语料库,带有UMLS语义类型标注,这些语言中立的编码提供了共享的跨语言标签空间,可直接与现有基于UMLS的资源对比。该语料库包含4,392篇PubMed摘要对、44,892个英越句子对及202,949个对齐的实体提及对,覆盖21种从MedMentions ST21pv数据集改编的语义类型。为平衡质量与可扩展性,采用自动翻译、专家后编辑、大模型辅助标签投影及人工验证与仲裁构建。将En-ViMedNER定为大规模银标语料库,并设有经人工审核和共识修正的微型测试子集。在两种场景下评估:(i) 越南语输入/输出生物医学NER;(ii) 英语输入/越南语输出跨语言NER。越南语NER基准测试了越南语监督编码器模型、英语监督多语言编码器模型及提示式大模型,最佳模型在测试集上获得52.70的F1,在微型测试集上达53.78。跨语言NER测试了编码器-解码器模型和提示式大模型,最佳模型在微型测试集上取得45.44的F1。我们公开发布语料库、构建流程及基线模型,以推动未来越南语生物医学自然语言处理研究。

原文摘要 · Abstract (English)

Biomedical Named Entity Recognition (NER) is fundamental to healthcare AI applications, including clinical decision support and medical information extraction. While corpora with Unified Medical Language System (UMLS) annotations, such as MedMentions, have driven progress in English biomedical NER, no comparable resource exists for Vietnamese. This paper presents En-ViMedNER, the first English-Vietnamese parallel biomedical NER corpus annotated with UMLS semantic types, which are language-neutral codes providing a shared cross-lingual label space and ensuring direct comparability with existing UMLS-based resources. The corpus contains 4,392 PubMed abstract pairs, 44,892 English-Vietnamese sentence pairs, and 202,949 aligned entity-mention pairs across 21 semantic types adapted from the MedMentions ST21pv dataset. To balance quality and scalability, we have constructed the corpus through automatic translation, expert post-editing, LLM-assisted label projection, and human verification and adjudication. We characterize En-ViMedNER as a large-scale silver-standard corpus with a human-audited and consensus-corrected mini-test subset. We evaluate En-ViMedNER in two settings: (i) Vietnamese-input/Vietnamese-output biomedical NER and (ii) English-input/Vietnamese-output cross-lingual NER. For Vietnamese NER, we benchmark Vietnamese-supervised encoder models, English-supervised multilingual encoder models, and prompt-based LLMs. The best model achieves an F1 score of 52.70 on the test set and 53.78 on the mini-test set. For cross-lingual NER, we benchmark encoder-decoder models and prompt-based LLMs. The best model achieves an F1 score of 45.44 on the mini-test set. We publicly release our corpus, corpus construction pipeline, and baseline models to facilitate future Vietnamese biomedical NLP research.

医学NER双语语料越南语NLPUMLS标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。