评估大模型生成的患者回复稿,看医生需要修改多少。
How Much Would a Clinician Edit This Draft? Evaluating LLM Alignment for Patient Message Response Drafting
- 构建医生回复的语义主题分类体系,量化修改负担。
- 模型在问诊引导类内容上修改量超70%,表现不佳。
- 针对医生风格定制提示词可显著降低修改需求,适合临床部署。
大型语言模型(LLMs)在生成患者门户消息回复方面展现潜力,但其在临床工作流中的应用仍存疑虑,例如是否真能节省医生时间。本文通过全面评估患者消息回复任务,研究了LLM与个体医生的对齐程度。我们提出一种新的回复主题分类体系,并构建评估框架,在内容和主题层面衡量医生对LLM生成回复的编辑负荷。我们发布了一个专家标注的数据集,对本地及商用LLM在多种适配技术(如主题提示、检索增强生成、监督微调、直接偏好优化)下的表现进行大规模评估。结果表明,模型在对齐医生回复时存在显著认知不确定性:虽能在部分主题上生成有效内容,但在需向患者追问信息的主题上表现差,编辑负荷超过70%。基于主题的适配策略在多数主题上带来改进。研究强调,必须针对个体医生偏好调整模型,才能实现可靠、负责任的医患沟通应用。
原文摘要 · Abstract (English)
Large language models (LLMs) show promise in drafting responses to patient portal messages, yet their integration into clinical workflows raises various concerns, including whether they would actually save clinicians time and effort in their portal workload. We investigate LLM alignment with individual clinicians through a comprehensive evaluation of the patient message response drafting task. We develop a novel taxonomy of thematic elements in clinician responses and propose a novel evaluation framework for assessing clinician editing load of LLM-drafted responses at both content and theme levels. We release an expert-annotated dataset and conduct large-scale evaluations of local and commercial LLMs using various adaptation techniques including thematic prompting, retrieval-augmented generation, supervised fine-tuning, and direct preference optimization. Our results reveal substantial epistemic uncertainty in aligning LLM drafts with clinician responses. While LLMs demonstrate capability in drafting certain thematic elements, they struggle with clinician-aligned generation in other themes, particularly question asking to elicit further information from patients. Theme-driven adaptation strategies yield improvements across most themes. Our findings underscore the necessity of adapting LLMs to individual clinician preferences to enable reliable and responsible use in patient-clinician communication workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。