测试AI医生沟通能力,发现协作重写最接近真人医生。
Can "AI" Be a Doctor? A Study of Empathy, Readability, and Alignment in Clinical LLMs

- 用多维度指标评估大模型医疗对话表现
- 协作重写使语义相似度达0.93,降低情感极端性
- 患者更喜欢重写版本,但专家仍更信真人医生
大型语言模型在医疗领域应用日益广泛,但其与临床沟通标准的契合度尚未充分量化。本研究对通用和专业领域大模型在结构化医学解释及真实医患互动中的表现进行多维度评估,分析语义准确性、可读性和情感共鸣。基线模型的情感极性比医生更强(极度负面:43.14-45.10% vs. 37.25%),且在GPT-5、Claude等大模型中语言复杂度显著更高(弗莱施-金凯德等级最高达16.91-17.60,医生为11.47-12.50)。情感导向提示可降低极端负面情绪并减少复杂度(GPT-5最多下降6.87分),但未提升语义准确度。协同重写效果最佳:语义相似度最高达0.93,同时改善可读性并减弱情感极端。双主体评估显示,无模型在知识性上超越医生,但患者一致偏好重写版本的清晰度与情感基调。结果表明,大模型最适合作为临床沟通的协作增强工具而非替代品。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly deployed in healthcare, yet their communicative alignment with clinical standards remains insufficiently quantified. We conduct a multidimensional evaluation of general-purpose and domain-specialized LLMs across structured medical explanations and real-world physician-patient interactions, analyzing semantic fidelity, readability, and affective resonance. Baseline models amplify affective polarity relative to physicians (Very Negative: 43.14-45.10% vs. 37.25%) and, in larger architectures such as GPT-5 and Claude, produce substantially higher linguistic complexity (FKGL up to 16.91-17.60 vs. 11.47-12.50 in physician-authored responses). Empathy-oriented prompting reduces extreme negativity and lowers grade-level complexity (up to -6.87 FKGL points for GPT-5) but does not significantly increase semantic fidelity. Collaborative rewriting yields the strongest overall alignment. Rephrase configurations achieve the highest semantic similarity to physician answers (up to mean = 0.93) while consistently improving readability and reducing affective extremity. Dual stakeholder evaluation shows that no model surpasses physicians on epistemic criteria, whereas patients consistently prefer rewritten variants for clarity and emotional tone. These findings suggest that LLMs function most effectively as collaborative communication enhancers rather than replacements for clinical expertise.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。