用强化学习让大模型更准识别中毒药物,提升急诊决策能力。
Learning Diagnostic Reasoning for Decision Support in Toxicology
- 用强化学习优化大模型,融合急救描述与生命体征数据。
- 在14类毒物上多标签预测微F1达0.644,超越专家水平。
- 适合急诊、中毒诊断等高风险医疗场景的智能辅助系统。
急性多物质中毒需在高度不确定下快速做出救命决策,临床医生常依赖不完整的摄入信息和非特异性症状。有效诊断推理需融合非结构化的非医学叙述(如急救人员现场描述、不可靠的患者自述或既往史)与结构化的医疗数据(如生命体征)。尽管大语言模型(LLMs)在处理此类异构输入方面具有潜力,但在此场景中表现不佳,常弱于仅依赖病史的简单基线。为此,我们提出DeToxR(带推理的中毒决策支持),首个将强化学习(RL)应用于急诊毒理学的模型。我们设计了一个鲁棒的数据融合引擎,基于经组相对策略优化(GRPO)微调的LLM,实现对14种物质类别的多标签预测。通过以临床性能为奖励信号直接优化模型推理过程,并将多标签一致度作为奖励,模型被明确惩罚遗漏共摄入物质或虚构不存在的毒物。实验表明,该模型显著优于未适配的基线LLM和监督基线。在临床验证中,其识别正确毒物的能力超越了一位专家毒理学家(微F1: 0.644 vs. 0.473)。结果证明,经过强化学习对齐的LLM能够有效整合非结构化院前叙述与结构化医疗数据,为高风险环境下的决策支持提供可能。
原文摘要 · Abstract (English)
Acute poly-substance intoxication requires rapid, life-saving decisions under substantial uncertainty, as clinicians must rely on incomplete ingestion details and nonspecific symptoms. Effective diagnostic reasoning in this chaotic environment requires fusing unstructured, non-medical narratives (e.g. paramedic scene descriptions and unreliable patient self-reports or known histories), with structured medical data like vital signs. While Large Language Models (LLMs) show potential for processing such heterogeneous inputs, they struggle in this setting, often underperforming simple baselines that rely solely on patient histories. To address this, we present DeToxR (Decision-support for Toxicology with Reasoning), the first adaptation of Reinforcement Learning (RL) to emergency toxicology. We design a robust data-fusion engine for multi-label prediction across 14 substance classes based on an LLM finetuned with Group Relative Policy Optimization (GRPO). We optimize the model's reasoning directly using a clinical performance reward. By formulating a multi-label agreement metric as the reward signal, the model is explicitly penalized for missing co-ingested substances and hallucinating absent poisons. Our model significantly outperforms its unadapted base LLM counterpart and supervised baselines. Furthermore, in a clinical validation study, the model indicates a clinical advantage by outperforming an expert toxicologist in identifying the correct poisons (Micro-F1: 0.644 vs. 0.473). These results demonstrate the potential of RL-aligned LLMs to synthesize unstructured pre-clinical narratives and structured medical data for decision support in high-stakes environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。