评测大模型在医疗对话中的社交能力,发现其表现参差不齐。
How sensitive do we want AI to be? Socio-communicative competencies of large language models in healthcare
- 用医疗咨询对话数据集,评估三款大模型的社交行为
- 非攻击性表现好,但结构化和共情能力弱
- 适合关注AI医疗交互设计的研究者与开发者
有效临床实践高度依赖医务人员的社交沟通能力。大语言模型(LLMs)被用于分诊、报告撰写或翻译医学术语等任务,以支持决策。这些应用既需事实准确性,也需社交胜任力。本研究通过分析来自HELP-Med数据集的1800条对话转录文本,评估GPT 4o、Llama 3和Command R+三款模型在与人类参与者互动时表现出的社会沟通能力。两名专家使用IC-MD量表对非敌意性、敏感性、结构化和非侵扰性进行编码。结果显示,模型在非敌意性方面表现良好,但在敏感性和非侵扰性上结果不一,结构化能力较差。结论指出,当前大模型尚缺乏稳定可靠的社交沟通技能,难以安全有效地作为医疗顾问。现有评估框架可助力开发更具社交响应性的模型,但需调整以适应人机行为差异。
原文摘要 · Abstract (English)
Background. Effective clinical practice relies heavily on the socio-communicative skills of medical professionals. Large language models (LLMs) have been proposed for tasks such as triaging patients, report drafting or translating medical jargon to support informed decision-making. These applications require both factual and social competence. This study evaluates dialogues between LLMs and participants to assess the current state of socio-communicative competencies displayed in LLM-generated texts. Methods. We extracted a subset of extended dialogues from the HELP-Med dataset, comprising 1800 conversation transcripts of interactions between human participants seeking medical information and three different LLMs, GPT 4o, Llama 3 and Command R+. Two experts coded the transcripts for demonstrations of socio-communicative behaviours (non-hostility, sensitivity, structuring, non-intrusiveness) using the IC-MD instrument, originally designed to evaluate interactional competencies in medical student admissions. Results. The LLMs in our study showed strength in non-hostility, mixed results in sensitivity and non-intrusiveness and performed poorly in structuring. Conclusion. Current LLMs lack the consistent and reliable socio-communicative skills needed for safe and effective use as healthcare advisors. While existing frameworks for assessing interactional competencies may support the development of more socially responsive LLMs, they will require adaptation to account for the differences in desirable behaviour between humans and LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。