arXiv:2502.01691cs.CLcs.AI2025-02被引 2

用智能体+不确定性感知提升医学报告结构化提取准确率

Agent-Based Uncertainty Awareness Improves Automated Radiology Report Labeling with an Open-Source Large Language Model

  • 设计智能体决策模型融合多提示输出,动态评估预测可信度
  • 过滤高不确定结果后F1提升至0.4787,卡帕系数达0.4258
  • 适用于非英文医学文本处理,提升医疗AI可解释性

使用大型语言模型(LLM)从放射科报告中可靠提取结构化数据仍具挑战性,尤其在复杂非英语文本(如希伯来语)中。本研究针对克罗恩病患者9,683份希伯来语放射科报告(2010–2023年,来自三家医疗机构)提出基于智能体的不确定性感知方法。512份报告经人工标注六项胃肠道器官与15种病理发现,其余采用HSMP-BERT自动标注。利用Llama 3-8b-instruct结合贝叶斯提示集成(BayesPE),通过六个语义等价提示估计不确定性。智能体决策模型将多提示输出整合为五级置信度,进行校准并对比三种基于熵的模型。评估指标包括准确率、F1分数、精确率、召回率及科恩卡帕系数。该模型在所有指标上均优于基线,原始F1为0.3967,召回率为0.6437,卡帕系数为0.3006;过滤高不确定性样本(≥0.5)后,F1提升至0.4787,卡帕增至0.4258。不确定性直方图显示正确与错误预测清晰分离,该模型提供最校准的不确定性估计。结合不确定性感知提示集与智能体决策机制,显著提升LLM在放射科报告结构化提取中的性能与可靠性,为高风险医疗应用提供更可解释、可信的解决方案。

原文摘要 · Abstract (English)

Reliable extraction of structured data from radiology reports using Large Language Models (LLMs) remains challenging, especially for complex, non-English texts like Hebrew. This study introduces an agent-based uncertainty-aware approach to improve the trustworthiness of LLM predictions in medical applications. We analyzed 9,683 Hebrew radiology reports from Crohn's disease patients (from 2010 to 2023) across three medical centers. A subset of 512 reports was manually annotated for six gastrointestinal organs and 15 pathological findings, while the remaining reports were automatically annotated using HSMP-BERT. Structured data extraction was performed using Llama 3.1 (Llama 3-8b-instruct) with Bayesian Prompt Ensembles (BayesPE), which employed six semantically equivalent prompts to estimate uncertainty. An Agent-Based Decision Model integrated multiple prompt outputs into five confidence levels for calibrated uncertainty and was compared against three entropy-based models. Performance was evaluated using accuracy, F1 score, precision, recall, and Cohen's Kappa before and after filtering high-uncertainty cases. The agent-based model outperformed the baseline across all metrics, achieving an F1 score of 0.3967, recall of 0.6437, and Cohen's Kappa of 0.3006. After filtering high-uncertainty cases (greater than or equal to 0.5), the F1 score improved to 0.4787, and Kappa increased to 0.4258. Uncertainty histograms demonstrated clear separation between correct and incorrect predictions, with the agent-based model providing the most well-calibrated uncertainty estimates. By incorporating uncertainty-aware prompt ensembles and an agent-based decision model, this approach enhances the performance and reliability of LLMs in structured data extraction from radiology reports, offering a more interpretable and trustworthy solution for high-stakes medical applications.

医学AI不确定性建模结构化提取智能体系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。