arXiv:2411.14487cs.CLcs.AI2024-11被引 15

提出医疗AI安全评估框架,发现主流大模型在医学任务中表现远逊于医生。

Ensuring Safety and Trust: Analyzing the Risks of Large Language Models in Medicine

  • 基于五项核心原则构建医疗AI安全评估体系
  • 11个大模型在1000道专家验证题中普遍表现不佳
  • 强调需人工监督与安全防护机制,适合医疗AI研发者参考

大型语言模型(LLMs)在医疗应用中展现出巨大潜力,但其潜在风险尚未系统性评估。本文提出五大医疗AI安全核心原则:真实性、鲁棒性、公平性、抗干扰性与隐私保护,并细化为十项具体维度。在此框架下,我们构建了包含1000道专家验证问题的MedGuard基准测试。对11种常用大模型的评估显示,当前模型无论是否经过安全对齐训练,在多数任务上表现均显著低于人类医师水平。尽管近期报告称先进模型如ChatGPT可在部分医疗任务中媲美甚至超越人类,本研究仍揭示出显著的安全差距,强调必须加强人工监督与部署AI安全防护机制。

原文摘要 · Abstract (English)

The remarkable capabilities of Large Language Models (LLMs) make them increasingly compelling for adoption in real-world healthcare applications. However, the risks associated with using LLMs in medical applications have not been systematically characterized. We propose using five key principles for safe and trustworthy medical AI: Truthfulness, Resilience, Fairness, Robustness, and Privacy, along with ten specific aspects. Under this comprehensive framework, we introduce a novel MedGuard benchmark with 1,000 expert-verified questions. Our evaluation of 11 commonly used LLMs shows that the current language models, regardless of their safety alignment mechanisms, generally perform poorly on most of our benchmarks, particularly when compared to the high performance of human physicians. Despite recent reports indicate that advanced LLMs like ChatGPT can match or even exceed human performance in various medical tasks, this study underscores a significant safety gap, highlighting the crucial need for human oversight and the implementation of AI safety guardrails.

医疗AI大模型安全评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。