arXiv:2510.27535cs.CL2025-10

让AI总结病历更懂患者需求,而非只看身体指标。

Patient-Centered Summarization Framework for AI Clinical Summarization: A Mixed-Methods Design

  • 基于患者与医生访谈设计新标准,强调价值观与生活背景
  • 5个开源模型在零样本/少样本下表现接近人类,但患者中心性仍逊于真人
  • 适合关注医疗AI伦理与可解释性的研究者和临床开发者

大型语言模型(LLMs)在生成患者-医生对话的临床摘要方面展现出接近人类水平的潜力。然而,现有摘要多聚焦于患者生理状况,忽视其偏好、价值观、愿望与担忧。为实现以患者为中心的照护,我们提出人工智能临床摘要的新标准:患者中心摘要(PCS)。目标是构建能捕捉患者价值观、兼具临床实用性的摘要框架,并评估当前开源大模型是否能在该任务中达到人类水平。采用混合方法研究:英国两组患者与公众参与(10名患者、8名医生),通过半结构化访谈确定摘要应包含的个人与情境信息及其结构。结果用于制定标注指南,由八名医生对88例心房颤动门诊记录生成金标准PCS。16例用于优化提示模板。五款开源模型(Llama-3.2-3B、Llama-3.1-8B、Mistral-8B、Gemma-3-4B、Qwen3-8B)对72例使用零样本与少样本提示生成摘要,评估指标包括ROUGE-L、BERTScore及定性分析。患者强调生活方式、社会支持、近期压力源与照护价值观;医生需要简洁的功能性、心理社会与情感背景信息。最佳零样本性能为Mistral-8B(ROUGE-L 0.189)和Llama-3.1-8B(BERTScore 0.673);最佳少样本为Llama-3.1-8B(ROUGE-L 0.206,BERTScore 0.683)。模型与专家在完整性与流畅性上相似,但正确性与患者中心性仍优于人类摘要。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly demonstrating the potential to reach human-level performance in generating clinical summaries from patient-clinician conversations. However, these summaries often focus on patients' biology rather than their preferences, values, wishes, and concerns. To achieve patient-centered care, we propose a new standard for Artificial Intelligence (AI) clinical summarization tasks: Patient-Centered Summaries (PCS). Our objective was to develop a framework to generate PCS that capture patient values and ensure clinical utility and to assess whether current open-source LLMs can achieve human-level performance in this task. We used a mixed-methods process. Two Patient and Public Involvement groups (10 patients and 8 clinicians) in the United Kingdom participated in semi-structured interviews exploring what personal and contextual information should be included in clinical summaries and how it should be structured for clinical use. Findings informed annotation guidelines used by eight clinicians to create gold-standard PCS from 88 atrial fibrillation consultations. Sixteen consultations were used to refine a prompt aligned with the guidelines. Five open-source LLMs (Llama-3.2-3B, Llama-3.1-8B, Mistral-8B, Gemma-3-4B, and Qwen3-8B) generated summaries for 72 consultations using zero-shot and few-shot prompting, evaluated with ROUGE-L, BERTScore, and qualitative metrics. Patients emphasized lifestyle routines, social support, recent stressors, and care values. Clinicians sought concise functional, psychosocial, and emotional context. The best zero-shot performance was achieved by Mistral-8B (ROUGE-L 0.189) and Llama-3.1-8B (BERTScore 0.673); the best few-shot by Llama-3.1-8B (ROUGE-L 0.206, BERTScore 0.683). Completeness and fluency were similar between experts and models, while correctness and patient-centeredness favored human PCS.

医疗AI患者中心摘要生成大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。