构建多语言临床病例数据集,助力低资源医学信息抽取。
Low-resource Information Extraction with the European Clinical Case Corpus
- 用大模型自动投影+人工校正,构建跨语言医疗文本数据集。
- 在5种语言上验证,微调后模型性能显著提升。
- 适合做多语言医疗信息抽取、低资源场景下的迁移学习研究。
我们提出E3C-3.0,一个涵盖五种源语(英语、法语、意大利语、西班牙语、巴斯克语)及五种目标语言(希腊语、意大利语、波兰语、斯洛伐克语、斯洛文尼亚语)的多语言医学临床病例数据集,包含疾病与检验结果关系标注。采用半自动方法,基于大语言模型(LLMs)进行自动标注投影,并辅以人工修订。实验表明,当前最先进的大模型在E3C-3.0上微调后表现更优;跨语言迁移学习效果显著,有效缓解数据稀缺问题。对比分析了原生语言数据与投影数据上的性能表现。数据已开源:https://huggingface.co/collections/NLP-FBK/e3c-projected-676a7d6221608d60e4e9fd89。
原文摘要 · Abstract (English)
We present E3C-3.0, a multilingual dataset in the medical domain, comprising clinical cases annotated with diseases and test-result relations. The dataset includes both native texts in five languages (English, French, Italian, Spanish and Basque) and texts translated and projected from the English source into five target languages (Greek, Italian, Polish, Slovak, and Slovenian). A semi-automatic approach has been implemented, including automatic annotation projection based on Large Language Models (LLMs) and human revision. We present several experiments showing that current state-of-the-art LLMs can benefit from being fine-tuned on the E3C-3.0 dataset. We also show that transfer learning in different languages is very effective, mitigating the scarcity of data. Finally, we compare performance both on native data and on projected data. We release the data at https://huggingface.co/collections/NLP-FBK/e3c-projected-676a7d6221608d60e4e9fd89 .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。