让大模型学会在癫痫治疗中推荐用药并适时放弃判断,提升基层医疗精准度。
Teaching LLMs to Recommend and Defer in Underrepresented Epilepsy Care

- 基于少量病历数据学习本地用药习惯,用提示记忆动态调整模型决策。
- 在乌干达两组独立数据上,用药推荐准确率比基线高4-8个百分点。
- 能自动识别高置信度病例,低置信度则主动推迟,适合资源有限地区使用。
在资源匮乏地区,专业癫痫诊疗能力稀缺,基于大语言模型的决策支持对一线医生管理长期治疗具有吸引力。这类系统需适应本地用药习惯,并懂得何时应推迟判断。本文聚焦乌干达儿科癫痫护理,从纵向非结构化门诊记录中预测抗癫痫药物方案。标准提示方法与医生处方达成非平凡一致性,但神经科医生评审发现,多数错误源于处方默认值与本地分布不匹配,而非理解本地病历失败。我们提出MANANA,一种非参数化提示学习框架,通过小规模患者级训练集学习本地用药指导。MANANA将观察到的处方错误转化为可审计的提示记忆,实现单代理和多代理变体,在两个独立收集的乌干达队列中均优于传统机器学习、直接提示及提示优化基线。进一步提出贝叶斯提示平均,将学习到的提示轨迹转化为处方概率和基于不确定性的推迟信号。在独立保留队列上,该方法使就诊级别前3名推荐准确率提升4-8个百分点,并支持选择性预测:系统可自动处理最自信的一半病例,精度达95%;或处理最自信的四分之一,精度达99%,其余低置信度案例则延迟至专科医生审查。
原文摘要 · Abstract (English)
Specialist epilepsy expertise is scarce in resource-constrained settings, making LLM-based decision support attractive for frontline clinicians managing longitudinal treatment. Such systems must adapt to local prescribing practice and know when to defer. We study this problem in Ugandan pediatric epilepsy care, predicting anti-seizure medication regimens from longitudinal unstructured clinic notes. Standard prompting achieves non-trivial agreement with physician prescriptions, but neurologist review shows that many errors reflect distribution-miscalibrated prescribing defaults rather than failures to parse the local record. We introduce MANANA, a non-parametric prompt-learning framework that learns local prescribing guidance from a small patient-level training set. MANANA converts observed prescription errors into auditable prompt memories, instantiated in single-agent and multi-agent variants, and improves over classical ML models, direct LLM prompting, and prompt-optimization baselines across two independently collected Ugandan cohorts. We further propose Bayesian prompt averaging, which converts the learned prompt trajectory into prescription likelihoods and an uncertainty-based deferral signal. On the independently collected held-out cohort, this improves visit-level top-3 prescription accuracy by 4-8 percentage points over prompt-optimization baselines and enables selective prediction: the system can auto-handle the most confident half of cases at 95% precision, or the most confident quarter at 99% precision, while deferring lower-confidence cases for specialist review.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。