arXiv:2409.15326cs.HCcs.AI2024-09被引 2

专为医生设计的AI助手比通用模型更可信、更实用。

Evaluating the Impact of a Specialized LLM on Physician Experience in Clinical Decision Support: A Comparison of Ask Avo and ChatGPT-4

  • 用专为医生优化的检索增强系统提升回答质量
  • 在信任度、可用性等五项指标上全面优于ChatGPT-4
  • 适合临床决策支持系统的开发与落地

大型语言模型(LLMs)在临床决策支持中的应用日益受到关注,但幻觉和缺乏明确来源引用等问题使其难以在临床环境中使用。本研究评估了AvoMD开发的专用模型Ask Avo,该模型结合了专有的语言模型增强检索(LMAR)系统、内置视觉引文提示及针对医生交互优化的提示工程,在模拟临床场景中与ChatGPT-4进行对比。62名参与者对8个来自不同专科指南的临床问题分别向两模型提问,每项回答按可信度、可操作性、相关性、全面性和友好格式评分(1–5分)。结果显示,Ask Avo在所有指标上均显著优于ChatGPT-4:可信度(4.52 vs. 3.34,p<0.001)、可操作性(4.41 vs. 3.19,p<0.001)、相关性(4.55 vs. 3.49,p<0.001)、全面性(4.50 vs. 3.37,p<0.001)和友好格式(4.52 vs. 3.60,p<0.001)。结果表明,针对临床需求设计的专用大模型能显著提升医生使用体验。

原文摘要 · Abstract (English)

The use of Large language models (LLMs) to augment clinical decision support systems is a topic with rapidly growing interest, but current shortcomings such as hallucinations and lack of clear source citations make them unreliable for use in the clinical environment. This study evaluates Ask Avo, an LLM-derived software by AvoMD that incorporates a proprietary Language Model Augmented Retrieval (LMAR) system, in-built visual citation cues, and prompt engineering designed for interactions with physicians, against ChatGPT-4 in end-user experience for physicians in a simulated clinical scenario environment. Eight clinical questions derived from medical guideline documents in various specialties were prompted to both models by 62 study participants, with each response rated on trustworthiness, actionability, relevancy, comprehensiveness, and friendly format from 1 to 5. Ask Avo significantly outperformed ChatGPT-4 in all criteria: trustworthiness (4.52 vs. 3.34, p<0.001), actionability (4.41 vs. 3.19, p<0.001), relevancy (4.55 vs. 3.49, p<0.001), comprehensiveness (4.50 vs. 3.37, p<0.001), and friendly format (4.52 vs. 3.60, p<0.001). Our findings suggest that specialized LLMs designed with the needs of clinicians in mind can offer substantial improvements in user experience over general-purpose LLMs. Ask Avo's evidence-based approach tailored to clinician needs shows promise in the adoption of LLM-augmented clinical decision support software.

临床决策大模型医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。