构建意大利急诊数据的134项临床表单标注集,探索大模型零样本填表可行性。
Toward Automatic Filling of Case Report Forms: A Case Study on Data from an Italian Emergency Department
- 基于意大利急诊真实病历,构建134项条目标注数据集
- 大模型在零样本下可完成填表,但倾向保守输出'未知'
- 适合医疗人工智能、临床数据自动化研究者参考
病例报告表(CRF)是临床研究中收集患者数据的核心工具。随着语言技术的发展,利用大语言模型(LLM)从临床笔记中自动填充CRF成为热点,但缺乏标注的CRF数据严重制约了该任务进展。为此,我们构建了一个来自意大利急诊科的真实临床笔记数据集,其中包含134个需填写的预定义条目,并进行了人工标注。本文分析了数据特征,定义了填表任务与评估指标,并使用开源前沿大模型进行了初步实验。结果表明:(i)可在零样本设置下对意大利语临床笔记进行CRF填表;(ii)LLM存在偏差,如倾向于回答“未知”,需加以校正。
原文摘要 · Abstract (English)
Case Report Forms (CRFs) collect data about patients and are at the core of well-established practices to conduct research in clinical settings. With the recent progress of language technologies, there is an increasing interest in automatic CRF-filling from clinical notes, mostly based on the use of Large Language Models (LLMs). However, there is a general scarcity of annotated CRF data, both for training and testing LLMs, which limits the progress on this task. As a step in the direction of providing such data, we present a new dataset of clinical notes from an Italian Emergency Department annotated with respect to a pre-defined CRF containing 134 items to be filled. We provide an analysis of the data, define the CRF-filling task and metric for its evaluation, and report on pilot experiments where we use an open-source state-of-the-art LLM to automatically execute the task. Results of the case-study show that (i) CRF-filling from real clinical notes in Italian can be approached in a zero-shot setting; (ii) LLMs' results are affected by biases (e.g., a cautious behaviour favours "unknown" answers), which need to be corrected.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。