arXiv:2603.19517cs.CVcs.LG2026-03

构建医疗照片理解统一基准,测试模型在真实场景下的医学图像推理能力。

ReXInTheWild: A Unified Benchmark for Medical Photograph Understanding

  • 基于484张真实医学文献照片,设计955道临床多选题
  • 顶尖模型最高准确率78%,医疗专用模型仅37%
  • 发现四类错误模式,为模型改进提供方向

日常普通相机拍摄的照片已广泛用于远程医疗和在线健康交流,但尚无全面的基准评估视觉语言模型对这类医学内容的理解能力。分析此类图像需融合精细自然图像理解与特定医学推理,这对通用和专用模型均构成挑战。我们提出ReXInTheWild,一个包含955道由临床医生验证的多选题的基准,覆盖484张来自生物医学文献的照片,涵盖七个临床主题。在该基准上评估时,领先多模态大模型表现差异显著:Gemini-3达78%准确率,Claude Opus 4.5为72%,GPT-5为68%,而医疗专用模型MedGemma仅37%。系统性误差分析揭示四类常见错误,从低级几何偏差到高级推理失败,需采用不同缓解策略。ReXInTheWild提供了自然图像理解与医学推理交汇处的高挑战性、临床真实基准。数据集已公开于HuggingFace。

原文摘要 · Abstract (English)

Everyday photographs taken with ordinary cameras are already widely used in telemedicine and other online health conversations, yet no comprehensive benchmark evaluates whether vision-language models can interpret their medical content. Analyzing these images requires both fine-grained natural image understanding and domain-specific medical reasoning, a combination that challenges both general-purpose and specialized models. We introduce ReXInTheWild, a benchmark of 955 clinician-verified multiple-choice questions spanning seven clinical topics across 484 photographs sourced from the biomedical literature. When evaluated on ReXInTheWild, leading multimodal large language models show substantial performance variation: Gemini-3 achieves 78% accuracy, followed by Claude Opus 4.5 (72%) and GPT-5 (68%), while the medical specialist model MedGemma reaches only 37%. A systematic error analysis also reveals four categories of common errors, ranging from low-level geometric errors to high-level reasoning failures and requiring different mitigation strategies. ReXInTheWild provides a challenging, clinically grounded benchmark at the intersection of natural image understanding and medical reasoning. The dataset is available on HuggingFace.

医疗视觉多模态基准评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。