arXiv:2506.11666cs.CLcs.AI2025-06ACL被引 8

将临床文本标注数据转为结构化病历表单,解决数据稀缺问题

Converting Annotated Clinical Cases into Structured Case Report Forms

  • 利用信息抽取标注数据半自动转换成结构化CRF格式
  • 英文零样本槽填充准确率达67.3%,意大利语59.7%
  • 适合医疗NLP、临床数据结构化方向的研究者使用

病例报告表(CRFs)在医学研究中广泛应用,可确保临床研究结果的准确性、可靠性和有效性。然而,公开且高质量标注的CRF数据集稀缺,制约了CRF槽填充系统的发展。为缓解这一问题,本文提出利用已有信息抽取任务标注数据,通过半自动方法将其转换为结构化CRF。该方法应用于英意双语的E3C数据集,构建出新高质量数据集。实验表明,基于闭源大模型的零样本槽填充在英文上达到67.3%、意大利语59.7%;而三类开源模型表现更差,说明从临床笔记自动填充CRF仍具挑战性。数据集已发布于https://huggingface.co/collections/NLP-FBK/e3c-to-crf-67b9844065460cbe42f80166。

原文摘要 · Abstract (English)

Case Report Forms (CRFs) are largely used in medical research as they ensure accuracy, reliability, and validity of results in clinical studies. However, publicly available, wellannotated CRF datasets are scarce, limiting the development of CRF slot filling systems able to fill in a CRF from clinical notes. To mitigate the scarcity of CRF datasets, we propose to take advantage of available datasets annotated for information extraction tasks and to convert them into structured CRFs. We present a semi-automatic conversion methodology, which has been applied to the E3C dataset in two languages (English and Italian), resulting in a new, high-quality dataset for CRF slot filling. Through several experiments on the created dataset, we report that slot filling achieves 59.7% for Italian and 67.3% for English on a closed Large Language Models (zero-shot) and worse performances on three families of open-source models, showing that filling CRFs is challenging even for recent state-of-the-art LLMs. We release the datest at https://huggingface.co/collections/NLP-FBK/e3c-to-crf-67b9844065460cbe42f80166

医疗NLPCRF构建槽填充多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。