真实患者对话多样,需建模情绪与沟通风格才能提升医疗聊天机器人效果
The complexities of patient-centred conversational artificial intelligence
- 构建分层模拟器,分别建模病情、情绪、策略与表达风格
- 真人评估显示模拟对话与真实对话几乎无法区分(准确率55%)
- 沟通风格显著影响分诊结果,适合临床部署的系统需适配多样性
面向消费者的医疗聊天机器人由大语言模型驱动,日益用于症状评估。然而,当前开发与评估多基于合作且表达清晰的模拟患者。我们分析了2,053段真实患者-聊天机器人对话,发现用户间沟通模式与情感表达差异显著。为此,我们开发了一种患者模拟器,可独立建模临床内容、情绪状态、对话策略与沟通风格。在15名人类评审员参与的图灵式真实性评估中,模拟对话与真实对话几乎无法区分,人类识别准确率为55%。通过五种不同患者人格,在1,164例医师评分案例中评估了四种LLM在紧急程度判断中的表现。结果表明,沟通风格会显著影响分诊结果。以理想化而非真实互动为设计基础的患者中心型对话人工智能,可能在实际部署中表现不佳,并加剧健康不平等。
原文摘要 · Abstract (English)
Consumer-facing health chatbots powered by large language models (LLMs) are increasingly used for symptom assessment. However, chatbot development and evaluation often rely on cooperative, articulate, simulated patients. We analysed 2,053 real patient-chatbot conversations and found that communication patterns and expression of emotions vary widely across users. We developed a patient simulator that separately models clinical content, emotional state, conversational strategy, and communication style. In a Turing-inspired evaluation of realism with 15 human graders, simulated conversations were nearly indistinguishable from real ones, with human graders achieving an accuracy of 55%. We used five distinct patient personae, across 1,164 clinician-graded cases, to evaluate the performance of four LLMs in urgency assessment. We found that communication style can significantly alter triage outcomes. Patient-centred conversational artificial intelligence must accommodate communication diversity: systems designed for idealised, rather than realistic, interactions risk underperforming and amplifying health disparities when deployed in the real world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。