用情境困境训练大模型,精准预测病人医疗偏好,效果超人类代理人。
Large Language Model Few-Shot Prompting with Dilemma Training Outperforms Human Surrogates in Predicting Patient Preferences

- 通过双向训练让模型在医疗困境中挖掘个体决策逻辑
- 准确率达81.7%,显著高于随机猜测和人类代理人的55%
- 适合用于设计能理解人性情境的智能医疗决策助手
在严重疾病情境下,人类代理人预测患者偏好准确率仅68%,易引发决策冲突。现有个性化偏好预测模型将价值视为静态评分,忽视医疗选择的情境依赖性。基于‘照护逻辑’,我们提出P4-DT(困境训练)模型,通过多样化医疗困境与用户互动,双向训练获取个体偏好推理过程。在12对患者-代理人配对的研究中,P4-DT预测患者治疗选择准确率达81.7%(优势比OR=5.61 [2.03, 15.51],p<.001),显著优于随机猜测及未辅助的代理人(55.0%;OR=3.67 [1.59, 8.47],p=.002),也优于受其辅助的代理人(61.7%)。对比提示分析显示,融入情境化决策与开放式文本可使准确率提升15.0个百分点。研究探讨了面向复杂决策的上下文感知型AI代理的未来发展方向。
原文摘要 · Abstract (English)
In serious illness, human surrogates often struggle to accurately predict patient preferences (68% accuracy), causing decision conflict. Personalized Patient Preference Predictor (P4) agents offer a potential solution, but prior prototypes treat values as static ratings, ignoring the contextual, situation-dependent nature of medical choices. Grounded in the 'logic of care', we present P4-DT (Dilemma Training), a P4 agent that constructs a patient decision policy by engaging users with varied medical dilemmas, eliciting individual preference reasoning through bi-directional training. In a study with 12 patient-surrogate dyads, P4-DT predicted patient treatment choices with 81.7% accuracy, significantly exceeding chance (OR = 5.61 [2.03, 15.51], p < .001) and outperforming both unassisted surrogates (55.0%; OR = 3.67 [1.59, 8.47], p = .002) and surrogates assisted by P4-DT (61.7%). Comparative prompt analyses showed that incorporating contextual scenario decisions and open-ended text improved accuracy by 15.0 percentage points over initial values ratings alone. We discuss implications for further testing and designing of context-aware AI agents that embody richer human experience to partner in complex decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。