arXiv:2511.18676cs.CVcs.AI2025-11中稿 · EMNLP

构建医学影像定量分析基准,提升模型测量肿瘤大小与角度的能力

MedVision: Benchmarking Quantitative Medical Image Analysis

  • 设计多模态医学影像数据集,支持结构检测、病灶尺寸估算与角度距离测量
  • 现有视觉语言模型在定量任务上表现差,微调后性能显著提升
  • 提出过程奖励与课程学习策略,助力模型实现精准医学量测

当前医学视觉语言模型主要面向分类问答(如“正常或异常”)或定性描述任务,但临床决策常依赖量化评估,如肿瘤大小或关节角度。此类定量推理能力在现有模型中仍严重不足。本文提出MedVision,一个大规模数据集与基准,专用于评估和提升模型在医学影像定量分析中的表现。该数据集涵盖22个公开数据集,包含29.0K 3D图像、11.2M标注2D切片及24.3M单实例标注。聚焦三类代表性定量任务:(1) 解剖结构与异常检测,(2) 肿瘤/病灶(T/L)尺寸估计,(3) 角度/距离(A/D)测量。实验表明,现有通用模型在此类任务上表现不佳。通过在MedVision上进行监督微调与强化学习微调(RFT),性能显著提升,得到MedVision-V0作为开源基线。RFT阶段引入过程奖励、乘法奖励组合及多任务课程学习策略并验证其有效性。MedVision为发展具备强定量推理能力的医学影像视觉语言模型提供了坚实基础。

原文摘要 · Abstract (English)

Current vision-language models (VLMs) in medicine are primarily designed for categorical question answering (e.g., "Is this normal or abnormal?") or qualitative descriptive tasks. However, clinical decision-making often relies on quantitative assessments, such as measuring the size of a tumor or the angle of a joint, from which clinicians draw their own diagnostic conclusions. This quantitative reasoning capability remains underexplored and poorly supported in existing VLMs. In this work, we introduce MedVision, a large-scale dataset and benchmark specifically designed to evaluate and improve VLMs on quantitative medical image analysis. MedVision spans 22 public datasets covering diverse anatomies and modalities, with 29.0K 3D images, 11.2M annotated 2D slices, and 24.3M single-instance annotations. We focus on three representative quantitative tasks: (1) detection of anatomical structures and abnormalities, (2) tumor/lesion (T/L) size estimation, and (3) angle/distance (A/D) measurement. We show that current off-the-shelf VLMs perform poorly on these tasks. However, supervised and reinforcement fine-tuning (RFT) on MedVision significantly enhances performance across detection, T/L size estimation, and A/D measurement, yielding MedVision-V0 as a strong open baseline. In the RFT stage, we design and evaluate the efficacy of process rewards, multiplicative reward composition, and multi-task RFT with curriculum learning. MedVision provides a foundation for developing VLMs with robust quantitative reasoning capabilities in medical imaging.

医学影像量化分析视觉语言模型基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。