arXiv:2509.18234cs.AIcs.CL2025-09被引 11

测试发现医疗AI在简单干扰下表现脆弱,真实应用仍需警惕。

The Illusion of Readiness in Health AI

  • 设计对抗性压力测试评估主流医疗AI模型鲁棒性。
  • 模型常在关键信息缺失时误判,且易被微小提示扰动误导。
  • 强调临床真实需求,呼吁超越排行榜的可信医疗AI标准。

大型语言模型在多项医学基准测试中展现出卓越性能,但在多模态推理等前沿领域仍存在显著短板。本文设计一系列对抗性压力测试,系统评估主流模型与医学基准的鲁棒性。结果显示,领先系统在关键输入缺失时仍能猜对答案,却极易受轻微提示变化干扰,且会生成看似合理实则错误的推理过程。通过临床专家制定的评分标准,我们发现现有医学基准测量目标差异巨大。研究揭示前沿AI在实际医疗应用中仍存在显著能力鸿沟。若要让AI赢得医疗信任,必须超越排行榜成绩,确保其具备鲁棒性、合理推理能力,并真正契合临床实际需求。

原文摘要 · Abstract (English)

Large language models have demonstrated remarkable performance in a wide range of medical benchmarks. Yet underneath the seemingly promising results lie salient growth areas, especially in cutting-edge frontiers such as multimodal reasoning. In this paper, we introduce a series of adversarial stress tests to systematically assess the robustness of flagship models and medical benchmarks. Our study reveals prevalent brittleness in the presence of simple adversarial transformations: leading systems can guess the right answer even with key inputs removed, yet may get confused by the slightest prompt alterations, while fabricating convincing yet flawed reasoning traces. Using clinician-guided rubrics, we demonstrate that popular medical benchmarks vary widely in what they truly measure. Our study reveals significant competency gaps of frontier AI in attaining real-world readiness for health applications. If we want AI to earn trust in healthcare, we must demand more than leaderboard wins and must hold AI systems accountable to ensure robustness, sound reasoning, and alignment with real medical demands.

医疗AI对抗测试鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。