GPT-4o在兽医病历信息提取中表现超人,准确率超96%。
Classification performance and reproducibility of GPT-4 omni for information extraction from veterinary electronic health records
- 用GPT-4o从兽医电子病历中自动提取临床症状
- 敏感度达96.9%,特异度97.6%,F1分数84.4%
- 错误多因病历歧义,适合自动化医疗数据处理
大型语言模型(LLMs)可从兽医电子健康记录(EHRs)中提取信息,但不同模型性能差异、温度设置影响及文本模糊性尚未评估。本研究比较了GPT-4 omni(GPT-4o)与GPT-3.5 Turbo在不同条件下的表现,并分析人类观察者一致性和LLM错误的关系。将LLMs与五名人类专家用于250份兽医转诊医院的EHR,识别与猫慢性肠病相关的六种临床体征。在温度0下,GPT-4o相对于人类多数意见,达到96.9%敏感度(四分位距[ IQR ] 92.9–99.3%)、97.6%特异度(IQR 96.5–98.5%)、80.7%阳性预测值(IQR 70.8–84.6%)、99.5%阴性预测值(IQR 99.0–99.9%)、84.4% F1分数(IQR 77.3–90.4%)和96.3%平衡准确率(IQR 95.0–97.9%)。GPT-4o性能显著优于其前代GPT-3.5 Turbo,尤其在敏感度上,后者仅达81.7%(IQR 78.9–84.8%)。调整GPT-4o温度未显著影响分类表现。无论温度如何,GPT-4o的可重复性均高于人类配对,温度0时平均Cohen's kappa为0.98(IQR 0.98–0.99),人类为0.8(IQR 0.78–0.81)。大多数错误出现在人类意见不一时(43例中的35例,81.4%),表明错误更可能源于病历模糊而非模型缺陷。使用GPT-4o自动提取兽医EHR信息是可行的替代方案。
原文摘要 · Abstract (English)
Large language models (LLMs) can extract information from veterinary electronic health records (EHRs), but performance differences between models, the effect of temperature settings, and the influence of text ambiguity have not been previously evaluated. This study addresses these gaps by comparing the performance of GPT-4 omni (GPT-4o) and GPT-3.5 Turbo under different conditions and investigating the relationship between human interobserver agreement and LLM errors. The LLMs and five humans were tasked with identifying six clinical signs associated with Feline chronic enteropathy in 250 EHRs from a veterinary referral hospital. At temperature 0, the performance of GPT-4o compared to the majority opinion of human respondents, achieved 96.9% sensitivity (interquartile range [IQR] 92.9-99.3%), 97.6% specificity (IQR 96.5-98.5%), 80.7% positive predictive value (IQR 70.8-84.6%), 99.5% negative predictive value (IQR 99.0-99.9%), 84.4% F1 score (IQR 77.3-90.4%), and 96.3% balanced accuracy (IQR 95.0-97.9%). The performance of GPT-4o was significantly better than that of its predecessor, GPT-3.5 Turbo, particularly with respect to sensitivity where GPT-3.5 Turbo only achieved 81.7% (IQR 78.9-84.8%). Adjusting the temperature for GPT-4o did not significantly impact classification performance. GPT-4o demonstrated greater reproducibility than human pairs regardless of temperature, with an average Cohen's kappa of 0.98 (IQR 0.98-0.99) at temperature 0 compared to 0.8 (IQR 0.78-0.81) for humans. Most GPT-4o errors occurred in instances where humans disagreed (35/43 errors, 81.4%), suggesting that these errors were more likely caused by ambiguity of the EHR than explicit model faults. Using GPT-4o to automate information extraction from veterinary EHRs is a viable alternative to manual extraction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。