用大模型模拟不同教育背景患者,评估医嘱理解度。
Enhancing Patient-Centric Communication: Leveraging LLMs to Simulate Patient Perspectives
- 用教育背景信息提示大模型,模拟患者对出院摘要的理解反应。
- 带背景提示时模型给出准确医疗建议率达88%,其他情况低于随机水平。
- 发现简单问答比定制化模拟更有效,适合医疗沟通优化研究。
大型语言模型(LLMs)在角色扮演任务中表现出色,尤其能通过特定提示模拟特定领域专家。本文评估了LLMs在模拟不同背景个体行为方面的有效性,并与真实人类反应进行对比。重点研究了模型对重症监护室(ICU)出院摘要的理解能力,分析了不同教育背景人群的可理解性差异。结果表明,当模型获得教育背景信息时,能以88%的准确率提供准确且可操作的医疗建议;但若仅提供其他信息,性能显著下降,低于随机水平。该初步研究揭示了自动生成多样化人群健康信息的潜力与风险,指出当前模型在临床应用中的关键差距。研究还发现,简单的问答模式可能优于复杂的人物定制策略,为个性化医疗沟通提供了重要起点。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated impressive capabilities in role-playing scenarios, particularly in simulating domain-specific experts using tailored prompts. This ability enables LLMs to adopt the persona of individuals with specific backgrounds, offering a cost-effective and efficient alternative to traditional, resource-intensive user studies. By mimicking human behavior, LLMs can anticipate responses based on concrete demographic or professional profiles. In this paper, we evaluate the effectiveness of LLMs in simulating individuals with diverse backgrounds and analyze the consistency of these simulated behaviors compared to real-world outcomes. In particular, we explore the potential of LLMs to interpret and respond to discharge summaries provided to patients leaving the Intensive Care Unit (ICU). We evaluate and compare with human responses the comprehensibility of discharge summaries among individuals with varying educational backgrounds, using this analysis to assess the strengths and limitations of LLM-driven simulations. Notably, when LLMs are primed with educational background information, they deliver accurate and actionable medical guidance 88% of the time. However, when other information is provided, performance significantly drops, falling below random chance levels. This preliminary study shows the potential benefits and pitfalls of automatically generating patient-specific health information from diverse populations. While LLMs show promise in simulating health personas, our results highlight critical gaps that must be addressed before they can be reliably used in clinical settings. Our findings suggest that a straightforward query-response model could outperform a more tailored approach in delivering health information. This is a crucial first step in understanding how LLMs can be optimized for personalized health communication while maintaining accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。