用大模型评估临终关怀沟通质量,效果优于传统方法。
PALLM: Evaluating and Enhancing PALLiative Care Conversations with Large Language Models

- 用医生标注的模拟对话测试大模型评估能力。
- 大模型在理解与共情等指标上表现优异,可生成带推理的反馈。
- 适合医疗AI研发者和临床系统设计者参考。
有效医患沟通对临床护理至关重要,直接影响患者结局与生活质量。传统评估方法如人工评分、患者反馈和医生自评存在成本高、难以扩展的问题。现有自然语言处理技术虽有潜力,但难以捕捉临床沟通细节,且需敏感临床数据训练,限制了实际应用。新兴大语言模型(LLMs)具备语言理解、上下文学习和推理能力,为评估复杂沟通指标提供了新路径,有望集成到被动监测与即时干预系统中。本研究探索将大模型用于评估姑息治疗沟通质量,利用医疗专业人员构建并标注的模拟对话脚本,测试专有模型(如GPT-4)及微调开源模型(如LLaMA2)的表现,以识别关键指标如‘理解’与‘共情’。结果表明,大模型在评估临床沟通方面表现更优,能提供带有推理过程的可操作反馈,并验证了自主研发内部大模型的可行性和实用性。该研究揭示了大模型提升医患互动的潜力,为构建大模型赋能的临床健康系统奠定了基础。
原文摘要 · Abstract (English)
Effective patient-provider communication is crucial in clinical care, directly impacting patient outcomes and quality of life. Traditional evaluation methods, such as human ratings, patient feedback, and provider self-assessments, are often limited by high costs and scalability issues. Although existing natural language processing (NLP) techniques show promise, they struggle with the nuances of clinical communication and require sensitive clinical data for training, reducing their effectiveness in real-world applications. Emerging large language models (LLMs) offer a new approach to assessing complex communication metrics, with the potential to advance the field through integration into passive sensing and just-in-time intervention systems. This study explores LLMs as evaluators of palliative care communication quality, leveraging their linguistic, in-context learning, and reasoning capabilities. Specifically, using simulated scripts crafted and labeled by healthcare professionals, we test proprietary models (e.g., GPT-4) and fine-tune open-source LLMs (e.g., LLaMA2) with a synthetic dataset generated by GPT-4 to evaluate clinical conversations, to identify key metrics such as `understanding' and `empathy'. Our findings demonstrated LLMs' superior performance in evaluating clinical communication, providing actionable feedback with reasoning, and demonstrating the feasibility and practical viability of developing in-house LLMs. This research highlights LLMs' potential to enhance patient-provider interactions and lays the groundwork for downstream steps in developing LLM-empowered clinical health systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。