arXiv:2506.13805cs.CYcs.AI2025-06被引 1

用真实用户健康问题测试大模型诊断能力,发现76%回复被医生认为准确。

Dr. GPT Will See You Now, but Should It? Exploring the Benefits and Harms of Large Language Models in Medical Diagnosis using Crowdsourced Clinical Cases

  • 通过众包收集212个真实健康疑问,评估大模型应对日常医疗咨询表现。
  • 平均76%的模型回复被9位认证医生判定为准确,部分回应存在潜在风险。
  • 引入RAG增强知识库后,回答质量提升,适合关注医疗AI应用的开发者和医生参考。

大型语言模型(LLMs)在高风险场景如医疗自诊与初步分诊中的应用日益广泛,引发对其有效性、适用性及潜在危害的伦理与实践担忧。现有研究多聚焦于专家设计的健康问题或医学考试题库,却忽视了对普通用户日常健康咨询的真实场景评估。为此,本文通过一场高校竞赛,采用新型众包方式,让34名参与者向4个公开可用的LLM提出212个真实或虚构的健康问题,由9位持证医师评估生成回复的准确性。结果显示,平均76%的回复被医生认为准确。同时,我们探讨了集成全面医学知识库的RAG版本是否能提升响应质量,并通过对7位医疗专业人士的访谈,提炼出解释量化结果的定性洞见。本研究旨在提供更贴近现实的医疗大模型应用表现认知。

原文摘要 · Abstract (English)

The proliferation of Large Language Models (LLMs) in high-stakes applications such as medical (self-)diagnosis and preliminary triage raises significant ethical and practical concerns about the effectiveness, appropriateness, and possible harmfulness of the use of these technologies for health-related concerns and queries. Some prior work has considered the effectiveness of LLMs in answering expert-written health queries/prompts, questions from medical examination banks, or queries based on pre-existing clinical cases. Unfortunately, these existing studies completely ignore an in-the-wild evaluation of the effectiveness of LLMs in answering everyday health concerns and queries typically asked by general users, which corresponds to the more prevalent use case for LLMs. To address this research gap, this paper presents the findings from a university-level competition that leveraged a novel, crowdsourced approach for evaluating the effectiveness of LLMs in answering everyday health queries. Over the course of a week, a total of 34 participants prompted four publicly accessible LLMs with 212 real (or imagined) health concerns, and the LLM generated responses were evaluated by a team of nine board-certified physicians. At a high level, our findings indicate that on average, 76% of the 212 LLM responses were deemed to be accurate by physicians. Further, with the help of medical professionals, we investigated whether RAG versions of these LLMs (powered with a comprehensive medical knowledge base) can improve the quality of responses generated by LLMs. Finally, we also derive qualitative insights to explain our quantitative findings by conducting interviews with seven medical professionals who were shown all the prompts in our competition. This paper aims to provide a more grounded understanding of how LLMs perform in real-world everyday health communication.

医疗AI大模型评估真实场景临床验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。