AI聊天机器人在文本交流中比医生更显共情,但需验证语音场景下的效果。
AI chatbots versus human healthcare professionals: a systematic review and meta-analysis of empathy in patient care
- 用大语言模型的AI聊天机器人与医生对比共情表现,基于13项研究合成分析
- AI平均得分高出0.87个标准差,相当于10分制上多出2分
- 适用于评估文本类医疗沟通中的共情表现,尤其关注生成式AI应用
背景:共情被广泛认为能改善患者结局,如减轻疼痛与焦虑、提升满意度,其缺失可能造成伤害。与此同时,基于人工智能(AI)的聊天机器人在医疗领域快速普及,五分之一的全科医生使用生成式AI辅助撰写信件等任务。部分研究显示AI聊天机器人在共情表现上优于人类医护人员(HCPs),但结果不一且缺乏整合。我们检索多个数据库,筛选出比较基于大语言模型的AI聊天机器人与人类医护人员在共情测量上的研究。采用ROBINS-I评估偏倚风险,可行时使用随机效应模型进行元分析,并避免重复计数。共识点:共识别15项研究(2023–2024年)。其中13项报告AI共情评分显著更高,仅2项皮肤病学研究倾向人类回应。15项中有13项数据可提取并纳入合并分析。对这13项均使用ChatGPT-3.5/4的研究进行元分析,结果显示标准化均值差为0.87(95% CI, 0.54–1.20),P < .00001,偏向AI,约相当于10分制上提升2分。争议点:研究依赖文本评估,忽略非语言线索,且共情由代理评分者判断。发展点:在纯文本场景下,AI聊天机器人常被感知为更具共情力。未来研究应通过患者直接评价验证该发现,并评估新型语音交互式AI是否也能实现类似共情优势。
原文摘要 · Abstract (English)
Background: Empathy is widely recognized for improving patient outcomes, including reduced pain and anxiety and improved satisfaction, and its absence can cause harm. Meanwhile, use of artificial intelligence (AI)-based chatbots in healthcare is rapidly expanding, with one in five general practitioners using generative AI to assist with tasks such as writing letters. Some studies suggest AI chatbots can outperform human healthcare professionals (HCPs) in empathy, though findings are mixed and lack synthesis. Sources of data: We searched multiple databases for studies comparing AI chatbots using large language models with human HCPs on empathy measures. We assessed risk of bias with ROBINS-I and synthesized findings using random-effects meta-analysis where feasible, whilst avoiding double counting. Areas of agreement: We identified 15 studies (2023-2024). Thirteen studies reported statistically significantly higher empathy ratings for AI, with only two studies situated in dermatology favouring human responses. Of the 15 studies, 13 provided extractable data and were suitable for pooling. Meta-analysis of those 13 studies, all utilising ChatGPT-3.5/4, showed a standardized mean difference of 0.87 (95% CI, 0.54-1.20) favouring AI (P < .00001), roughly equivalent to a two-point increase on a 10-point scale. Areas of controversy: Studies relied on text-based assessments that overlook non-verbal cues and evaluated empathy through proxy raters. Growing points: Our findings indicate that, in text-only scenarios, AI chatbots are frequently perceived as more empathic than human HCPs. Areas timely for developing research: Future research should validate these findings with direct patient evaluations and assess whether emerging voice-enabled AI systems can deliver similar empathic advantages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。