用视觉语言模型统一评估与修复医学影像,提升跨模态修复效果。
MedQ-UNI: Toward Unified Medical Image Quality Assessment and Restoration via Vision-Language Modeling
- 先评估后修复:通过自然语言描述质量问题,指导图像修复。
- 单模型在5种任务上达到顶尖效果,50K数据训练,2K样本评测。
- 适合需要通用、可解释医学图像修复的临床研究者。
现有医学图像修复方法通常针对特定模态或退化类型,难以泛化到临床中多样的退化情况。我们提出MedQ-UNI,一个基于视觉-语言建模的统一模型,采用‘评估-修复’范式,利用医学图像质量评估(Med-IQA)引导跨模态、跨退化类型的修复。该模型采用多模态自回归双专家架构,共享注意力机制:质量评估专家通过结构化自然语言描述识别退化问题,修复专家则据此生成针对性修复结果。为此,我们构建了一个约5万对样本的大规模数据集,涵盖三种成像模态和五类修复任务,每条样本均配有结构化质量描述,用于联合训练;同时建立2000样本基准测试集。大量实验表明,无需任务特定微调,单个MedQ-UNI模型在所有任务上均达当前最优性能,并生成更优的质量描述,证明显式质量理解显著提升修复精度与可解释性。
原文摘要 · Abstract (English)
Existing medical image restoration (Med-IR) methods are typically modality-specific or degradation-specific, failing to generalize across the heterogeneous degradations encountered in clinical practice. We argue this limitation stems from the isolation of Med-IR from medical image quality assessment (Med-IQA), as restoration models without explicit quality understanding struggle to adapt to diverse degradation types across modalities. To address these challenges, we propose MedQ-UNI, a unified vision-language model that follows an assess-then-restore paradigm, explicitly leveraging Med-IQA to guide Med-IR across arbitrary modalities and degradation types. MedQ-UNI adopts a multimodal autoregressive dual-expert architecture with shared attention: a quality assessment expert first identifies degradation issues through structured natural language descriptions, and a restoration expert then conditions on these descriptions to perform targeted image restoration. To support this paradigm, we construct a large-scale dataset of approximately 50K paired samples spanning three imaging modalities and five restoration tasks, each annotated with structured quality descriptions for joint Med-IQA and Med-IR training, along with a 2K-sample benchmark for evaluation. Extensive experiments demonstrate that a single MedQ-UNI model, without any task-specific adaptation, achieves state-of-the-art restoration performance across all tasks while generating superior descriptions, confirming that explicit quality understanding meaningfully improves restoration fidelity and interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。