arXiv:2512.01191cs.CL2025-12被引 7

通用大模型在医学任务上表现优于专业临床工具。

Generalist Large Language Models Outperform Clinical Tools on Medical Benchmarks

  • 用医学知识与临床对齐任务测试通用和专业模型
  • 通用模型平均得分高于专业工具,最高达GPT-5
  • 专业工具存在信息不全、沟通质量差等问题

专用临床AI助手正快速进入医疗实践,常被宣传为比通用大语言模型更安全可靠。然而,与前沿模型不同,这些临床工具很少接受独立的定量评估,导致其在诊断、分诊和指南解读中的影响缺乏证据支持。我们通过包含1000个样本的微型基准测试,评估了两种广泛使用的临床AI系统(OpenEvidence和UpToDate Expert AI)与三种最先进的通用大模型(GPT-5、Gemini 3 Pro、Claude Sonnet 4.5),任务涵盖MedQA(医学知识)和HealthBench( clinician-alignment)。结果显示,通用模型始终表现更优,其中GPT-5得分最高;而OpenEvidence和UpToDate在信息完整性、沟通质量、上下文理解及系统性安全推理方面存在明显缺陷。研究揭示,被标榜用于临床决策支持的工具可能普遍落后于前沿大模型,凸显了在患者应用场景部署前开展透明、独立评估的紧迫性。

原文摘要 · Abstract (English)

Specialized clinical AI assistants are rapidly entering medical practice, often framed as safer or more reliable than general-purpose large language models (LLMs). Yet, unlike frontier models, these clinical tools are rarely subjected to independent, quantitative evaluation, creating a critical evidence gap despite their growing influence on diagnosis, triage, and guideline interpretation. We assessed two widely deployed clinical AI systems (OpenEvidence and UpToDate Expert AI) against three state-of-the-art generalist LLMs (GPT-5, Gemini 3 Pro, and Claude Sonnet 4.5) using a 1,000-item mini-benchmark combining MedQA (medical knowledge) and HealthBench (clinician-alignment) tasks. Generalist models consistently outperformed clinical tools, with GPT-5 achieving the highest scores, while OpenEvidence and UpToDate demonstrated deficits in completeness, communication quality, context awareness, and systems-based safety reasoning. These findings reveal that tools marketed for clinical decision support may often lag behind frontier LLMs, underscoring the urgent need for transparent, independent evaluation before deployment in patient-facing workflows.

大模型医学AI评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。