arXiv:2607.04694cs.CV2026-07

让视觉语言模型学会处理原始杂乱医疗数据,解决真实临床应用中的关键难题。

Solve the Missing First Step: Can VLMs Standardize Raw Heterogeneous Medical Data?

论文配图:Solve the Missing First Step: Can VLMs Standardize Raw Heterogeneous Medical Data?
图 1 · 摘自论文原文
  • 设计新评测基准,测试模型从原始文件夹中自动识别并标准化医学图像与文本
  • 顶尖模型Gemini 3 Flash仅48.6%成功率,暴露数据预处理瓶颈
  • 适用于医疗AI落地研究者、多模态模型开发者及临床数据工程师

随着视觉语言模型(VLMs)在医疗AI中的广泛应用,现有评测主要聚焦于已标准化图像和文本上的诊断能力,隐含假设数据已准备好。然而在真实临床场景中,医疗数据常为原始、异构且分散在不同来源。本文研究这一被忽视的步骤——原始医疗数据标准化。具体而言,模型需处理原始数据文件夹,识别来源格式,将原始医学图像转换为VLM可用的视觉输入,提取相关文本信息,并组织成结构化图文对。为此,我们构建了医学数据标准化基准MDS-Bench,人工标注1,939个任务,覆盖多样临床实践、影像模态、标注格式与目录结构。大量实验表明,即使最强模型Gemini 3 Flash,端到端成功率为48.6%。研究揭示:原始数据标准化是真实医疗AI诊断中的关键瓶颈。

原文摘要 · Abstract (English)

As vision-language models (VLMs) are increasingly applied to medical AI, existing benchmarks mainly focus on evaluating their diagnosis ability over given medical images and texts, implicitly assuming that standardized medical images, texts or question-answer pairs are already prepared. However, this assumption does not hold when we apply VLMs in real clinical practice, where medical data is often raw, heterogeneous, and fragmented across different sources. In this paper, we study this missing step, i.e., raw medical data standardization. Specifically, models are given raw dataset folders and evaluated on their ability to identify source formats, convert raw medical images into VLM-compatible visual inputs, extract relevant textual information, and organize the results into structured image-text pairs. To construct this Medical Data Standardization Benchmark (MDS-Bench), we manually annotate 1,939 raw medical data standardization tasks covering diverse clinical practice, radiology modalities, annotation formats, and directory layouts. Extensive experiments show that even the best performing VLMs, i.e., Gemini 3 Flash, achieve only 48.6% end-to-end success rate. Our research highlights raw medical data standardization as a critical bottleneck for medical AI diagnosis in real practice.

视觉语言模型医疗数据标准化多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。