arXiv:2512.04238cs.CV2025-12被引 3

罕见解剖变异暴露视觉语言模型在医学影像中的致命弱点

6 Fingers, 1 Kidney: Natural Adversarial Medical Images Reveal Critical Weaknesses of Vision-Language Models

  • 构建首个自然发生的罕见解剖变异基准,涵盖多种成像模态
  • 模型准确率从正常情况的71%降至异常情况的28%,顶尖模型仍下降41%-51%
  • 现有方法无法缓解偏差,适合医疗AI安全评估与模型改进研究者

视觉语言模型(VLMs)正日益融入临床工作流程。然而,现有基准主要评估常见解剖表现下的性能,未能捕捉罕见变异带来的挑战。为此,我们提出AdversarialAnatomyBench,首个包含跨多种成像模态和解剖区域的自然罕见解剖变异基准。我们将违反对“典型”人体解剖先验的变异称为自然对抗性解剖。对25个前沿VLM进行基准测试,得出三项关键发现:第一,面对基础医学感知任务,平均准确率从典型解剖的71%降至非典型解剖的28%;即使最佳模型(GPT-5、Gemini 2.5 Pro、Llama 4 Maverick)也出现41%-51%的性能下降;第二,模型错误紧密对应预期解剖偏见;第三,模型扩展或干预措施(包括偏见感知提示与测试时推理)均无法解决此问题。这些结果揭示当前VLM在罕见解剖表现上的泛化能力存在严重缺陷。AdversarialAnatomyBench为系统测量和缓解多模态医疗人工智能系统的解剖偏见提供了基础。

原文摘要 · Abstract (English)

Vision-language models (VLMs) are increasingly integrated into clinical workflows. However, existing benchmarks primarily assess performance on common anatomical presentations and fail to capture the challenges posed by rare variants. To address this gap, we introduce AdversarialAnatomyBench, the first benchmark comprising naturally occurring rare anatomical variants across diverse imaging modalities and anatomical regions. We call such variants that violate learned priors about "typical" human anatomy natural adversarial anatomy. Benchmarking 25 state-of-the-art VLMs with AdversarialAnatomyBench yielded three key insights. First, when queried with basic medical perception tasks, mean accuracy dropped from 71% on typical to 28% on atypical anatomy. Even the best-performing models, GPT-5, Gemini 2.5 Pro, and Llama 4 Maverick, showed performance drops of 41-51%. Second, model errors closely mirrored expected anatomical biases. Third, neither model scaling nor interventions, including bias-aware prompting and test-time reasoning, resolved these issues. These findings highlight a critical limitation in current VLMs: their poor generalization to rare anatomical presentations. AdversarialAnatomyBench provides a foundation for systematically measuring and mitigating anatomical bias in multimodal medical artificial intelligence (AI) systems.

视觉语言模型医学影像罕见变异模型偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。