arXiv:2607.20428cs.CLcs.HC2026-07

用大模型辅助医生更快更准地识别皮肤免疫副作用

Human-in-the-Loop Large Language Model Framework for Identification of Cutaneous Immune-Related Adverse Events

  • 构建人机协作的多智能体大模型框架,结合检索增强技术
  • 检测准确率提升至F1=0.88,评阅时间减半,一致性显著提高
  • 适合临床研究、药物安全监测人员快速提取不良反应数据

本研究评估了一种检索增强的多智能体大语言模型(LLM)驱动的人机协同框架,用于从临床记录中识别皮肤免疫相关不良事件(cirAEs)。与人工独立审阅相比,该框架在准确率(F1=0.88 vs 0.77)、评分者间一致性(Cohen's kappa=0.82 vs 0.50)方面均有提升,并将平均审阅时间减少约一半。该框架为大模型在跨器官系统免疫毒性识别中的应用提供了范例,更广泛地支持了不良事件数据的精准、可扩展且透明的提取。

原文摘要 · Abstract (English)

This study evaluated a retrieval-augmented, multi-agent large language model (LLM)-driven, human-in-the-loop framework for detecting cutaneous immune-related adverse events (cirAEs) from clinical notes. Compared with unassisted manual review, the LLM-assisted workflow improved accuracy (F1 = 0.88 vs 0.77), inter-rater agreement measured by Cohen's kappa (kappa = 0.82 vs 0.50), and reduced average review time by approximately half. This framework pilots how LLMs can be applied to identify immune-related toxicities across organ systems and, more broadly, enable accurate, scalable, and transparent adverse event data extraction.

医学AI大模型不良反应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。