arXiv:2511.13135cs.CV2025-11被引 1

构建首个面向医疗多模态生成的上下文关联型评测基准

MedGEN-Bench: Contextually entangled benchmark for open-ended multimodal medical generation

  • 设计三类任务格式,要求跨模态推理与开放式生成
  • 包含6422对专家验证的图文对,覆盖16项临床任务
  • 引入像素级+语义分析+专家评分的三层评估体系

随着视觉语言模型在医疗领域应用日益广泛,临床医生期望AI不仅能生成文本诊断,还能生成可融入真实诊疗流程的医学影像。现有医疗图像评测基准存在明显缺陷:查询模糊、诊断推理简化为封闭式选择、评价偏重文本而忽视图像生成能力。为此,我们提出MedGEN-Bench,一个综合性多模态评测基准,包含6422对专家验证的图像-文本对,覆盖六种影像模态、16项临床任务和28个子任务。该基准分为视觉问答、图像编辑和上下文多模态生成三类格式,强调需复杂跨模态推理与开放式生成的上下文纠缠指令。我们设计了三层次评估框架,融合像素级指标、语义文本分析与专家指导的临床相关性评分,系统评估了10个组合框架、3个统一模型和5个VLMs。

原文摘要 · Abstract (English)

As Vision-Language Models (VLMs) increasingly gain traction in medical applications, clinicians are progressively expecting AI systems not only to generate textual diagnoses but also to produce corresponding medical images that integrate seamlessly into authentic clinical workflows. Despite the growing interest, existing medical visual benchmarks present notable limitations. They often rely on ambiguous queries that lack sufficient relevance to image content, oversimplify complex diagnostic reasoning into closed-ended shortcuts, and adopt a text-centric evaluation paradigm that overlooks the importance of image generation capabilities. To address these challenges, we introduce MedGEN-Bench, a comprehensive multimodal benchmark designed to advance medical AI research. MedGEN-Bench comprises 6,422 expert-validated image-text pairs spanning six imaging modalities, 16 clinical tasks, and 28 subtasks. It is structured into three distinct formats: Visual Question Answering, Image Editing, and Contextual Multimodal Generation. What sets MedGEN-Bench apart is its focus on contextually intertwined instructions that necessitate sophisticated cross-modal reasoning and open-ended generative outputs, moving beyond the constraints of multiple-choice formats. To evaluate the performance of existing systems, we employ a novel three-tier assessment framework that integrates pixel-level metrics, semantic text analysis, and expert-guided clinical relevance scoring. Using this framework, we systematically assess 10 compositional frameworks, 3 unified models, and 5 VLMs.

多模态生成医疗AI评测基准视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。