arXiv:2510.01691cs.CV2025-10被引 11

构建医学影像质量评估新基准,用大模型模拟医生判读逻辑。

MedQ-Bench: Evaluating and Exploring Medical Image Quality Assessment Abilities in MLLMs

  • 设计感知与推理双任务,让大模型像医生一样描述影像质量。
  • 覆盖5种成像模态、40余项属性,含真实临床与生成图像数据。
  • 发现当前大模型评估能力不稳定,需针对性优化以用于临床。

医学影像质量评估(IQA)是临床AI的首道安全关口,但现有方法受限于单一评分指标,无法体现专家评估中描述性、类人推理的核心特征。为此,我们提出MedQ-Bench,一个面向多模态大语言模型(MLLMs)的医学影像质量评估综合基准,建立基于语言的感知-推理范式。该基准包含两个互补任务:(1) MedQ-Perception,通过人工标注的问题测试模型对基础视觉属性的感知能力;(2) MedQ-Reasoning,涵盖无参考和对比推理任务,使模型评估贴近人类判读逻辑。基准覆盖五种成像模态及超过四十项质量属性,共包含2,600个感知查询与708个推理评估,数据来源包括真实临床图像、基于物理重建的模拟退化图像以及AI生成图像。为评估推理能力,我们提出多维评判协议,从四个互补维度评估模型输出。通过对比大模型判断与放射科医生意见,开展严格的医-智对齐验证。对14个前沿MLLMs的评估显示,模型具备初步但不稳定的感知与推理能力,准确率尚不足以支持可靠临床应用。研究结果凸显了在医学IQA中对MLLMs进行定向优化的必要性。我们期望MedQ-Bench能推动后续探索,释放MLLMs在医学影像质量评估中的潜力。

原文摘要 · Abstract (English)

Medical Image Quality Assessment (IQA) serves as the first-mile safety gate for clinical AI, yet existing approaches remain constrained by scalar, score-based metrics and fail to reflect the descriptive, human-like reasoning process central to expert evaluation. To address this gap, we introduce MedQ-Bench, a comprehensive benchmark that establishes a perception-reasoning paradigm for language-based evaluation of medical image quality with Multi-modal Large Language Models (MLLMs). MedQ-Bench defines two complementary tasks: (1) MedQ-Perception, which probes low-level perceptual capability via human-curated questions on fundamental visual attributes; and (2) MedQ-Reasoning, encompassing both no-reference and comparison reasoning tasks, aligning model evaluation with human-like reasoning on image quality. The benchmark spans five imaging modalities and over forty quality attributes, totaling 2,600 perceptual queries and 708 reasoning assessments, covering diverse image sources including authentic clinical acquisitions, images with simulated degradations via physics-based reconstructions, and AI-generated images. To evaluate reasoning ability, we propose a multi-dimensional judging protocol that assesses model outputs along four complementary axes. We further conduct rigorous human-AI alignment validation by comparing LLM-based judgement with radiologists. Our evaluation of 14 state-of-the-art MLLMs demonstrates that models exhibit preliminary but unstable perceptual and reasoning skills, with insufficient accuracy for reliable clinical use. These findings highlight the need for targeted optimization of MLLMs in medical IQA. We hope that MedQ-Bench will catalyze further exploration and unlock the untapped potential of MLLMs for medical image quality evaluation.

医学影像大模型质量评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。