用自动化流程从真实病历生成3D肿瘤影像问答数据集,评估视觉语言模型真能力。
Automated Report-Derived Oncology VQA Benchmark for Evaluating Vision-Language Models on 3D Medical Imaging

- 通过智能代理自动从病历与影像配对数据生成两类问题:标准报告式和基于大模型提炼的临床发现题。
- 在4个院内癌症队列上构建无污染基准,零样本测试显示所有模型仍有巨大提升空间。
- 揭示图像依赖性因数据集而异,部分任务闭眼反而表现更好,说明训练数据可能已泄露。
评估视觉语言模型(VLMs)在医学影像上的表现需要具有临床基础、可扩展且能控制混淆因素的基准。现有公开基准存在规模小、人工标注或可能泄露至VLM预训练语料等问题。本文提出一种自动化代理驱动的流水线,直接从私有放射科报告与3D肿瘤影像配对数据中生成多选题问答数据集,包含两类互补问题:基于临床报告模板的RADS风格问题,以及由大语言模型从放射科报告中提炼并验证的报告衍生问题。该方法应用于四个院内癌症队列,生成了无实例污染、无需逐题人工标注的基准。对六种VLM进行零样本评估,未出现主导模型,各维度均有显著提升空间。盲测消融实验表明,视觉依赖性高度依赖数据集:肝部报告类问题确实需图像支持,而肺部CT问题几乎可脱离图像解答——领先闭源模型在肺部CT上闭眼表现甚至优于睁眼,暗示即使私有临床数据也无法保证视觉能力评估的纯净性。该流水线作为开源代理技能发布,支持院内重部署。
原文摘要 · Abstract (English)
Evaluating vision-language models (VLMs) on medical images requires benchmarks that are clinically grounded, scalable, and controlled for evaluation confounds. Existing public benchmarks are limited in scale, manually annotated, or potentially leaked into VLM pretraining corpora. We present an automated agent-driven pipeline that generates multiple-choice VQA datasets directly from paired private radiology reports and 3D oncology imaging, producing two complementary question types: RADS-style questions deterministically derived from clinician-defined reporting schemas, and radiology report-derived questions generated by an LLM from radiologist findings and verified against the source report. Applied to four in-house cancer cohorts, the pipeline yields an instance-contamination-controlled benchmark without per-question human annotation. Zero-shot evaluation of six VLMs reveals no dominant model and substantial headroom across all cells. A blind ablation reveals that visual reliance is highly dataset-specific: liver Report-derived questions genuinely require the image, while Lung CT is essentially solvable without it - the leading closed model exceeds its sighted accuracy on Lung CT when blinded - indicating that even private clinical data does not guarantee a contamination-controlled read of visual capability. The pipeline is released as an open agent skill for in-house redeployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。