arXiv:2512.01241cs.CYcs.AI2025-12

医学AI助诊存在严重安全风险,需系统评估与改进。

First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations

  • 构建1100个临床案例的NOHARM基准,评估AI建议潜在危害。
  • 24.6%案例中直接使用AI建议可能造成严重伤害,80%为遗漏错误。
  • 人机协作可提升表现,但常忽略AI建议,潜力未被释放。

大型语言模型(LLMs)和医疗AI工具被医生与患者广泛用于获取医疗建议,但其临床安全性仍缺乏充分评估。本文提出NOHARM(Numerous Options Harm Assessment for Risk in Medicine),一个包含1,100个从初级保健到专科会诊案例的基准,用于衡量LLM生成的会诊建议中潜在有害错误的发生频率与严重程度。NOHARM涵盖10个专科领域,包含4,249个临床管理选项的12,747个专家标注。在20个主流LLMs和4个常用检索增强生成(RAG)临床AI工具中,直接应用建议可能导致严重伤害的案例达24.6%,其中超过80%为遗漏型错误。不同系统表现不一,临床专用AI优于通用大模型,多智能体协作进一步提升了通用模型表现。对101名美国注册全科医生的随机对照研究显示,相比传统资源,AI辅助能提升医生表现;然而,医生常忽略有价值的AI建议,整体得分仍低于多数独立AI系统。若采纳全部建议,人机联合响应将超越单独的人或AI,表明两者具有互补优势且协同潜力尚未实现。结果表明,尽管当前AI在医学知识评测中表现优异,其生成的会诊建议仍可能带来严重风险,亟需建立明确的临床安全评估机制。该基准与排行榜已公开,以支持临床AI系统的持续评估与优化。

原文摘要 · Abstract (English)

Large language models (LLMs) and medical AI tools are routinely used by physicians and patients for medical advice, yet their clinical safety profiles remain poorly characterized. We present NOHARM (Numerous Options Harm Assessment for Risk in Medicine), a 1,100-task benchmark of primary care-to-specialist consultation cases to measure the frequency and severity of potentially harmful errors from LLM-generated medical consultation recommendations. NOHARM covers 10 specialties, with 12,747 expert annotations for 4,249 clinical management options. Across 20 notable LLMs and 4 widely used retrieval-augmented generation (RAG) clinical AI tools, direct application of recommendations carried potential for severe harm in up to 24.6% of cases, with errors of omission accounting for more than 80% of severe errors. Harm potential was not uniform across systems, with clinical AI tools outperforming generalist LLMs, and multi-agent AI teaming further improving performance in generalist models. In a randomized study of 101 U.S.-licensed generalist physicians, AI assistance improved physician performance compared to conventional resources. However, AI-assisted physicians frequently omitted valuable AI-generated recommendations and still scored lower than many AI systems alone. Had those recommendations been incorporated, combined human-AI responses would have outperformed both the human and AI system as used, suggesting complementary strengths and unrealized potential in human-AI teaming. Collectively, these results show that despite strong performance on medical knowledge benchmarks, widely used AI tools can produce medical consultation advice with the potential for severe harm, and highlight the need for explicit measurement of clinical safety. The benchmark and leaderboard are publicly available to support ongoing evaluation and improvement of AI systems used for clinical care.

医疗AI安全评估人机协作大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。