arXiv:2509.24888cs.CVcs.CL2025-09

用多模态大模型提升MRI图像质量评估,兼顾准确与可解释性。

MMRQA: Signal-Enhanced Multimodal Large Language Models for MRI Quality Assessment

  • 融合信号处理与大模型,将量化指标转为问答对进行分析
  • 在三个数据集上达到顶尖性能,零样本泛化能力出色
  • 适合临床医疗场景,输出结果可读性强,便于医生理解

磁共振成像(MRI)质量评估对临床决策至关重要,但受限于数据稀缺和扫描协议差异。传统方法存在根本权衡:基于信号的方法(如MRIQC)提供量化指标却缺乏语义理解;深度学习方法虽精度高,却牺牲可解释性。为此,我们提出多模态MRI质量评估框架MMRQA,首次将多模态大语言模型(MLLMs)与采集感知信号处理相结合。MMRQA包含三项创新:通过引入模拟伪影的MRQy实现鲁棒指标提取,利用Qwen将指标结构化为问答对,采用低秩适配(LoRA)微调LLaVA-OneVision实现参数高效融合。在MR-ART、FastMRI和MyConnectome基准上评估,MMRQA达到当前最优表现,并通过全面消融实验验证了其零样本泛化能力。该框架实现了定量分析与语义推理的融合,生成可被临床理解的输出,显著提升动态医疗环境下的质量控制水平。

原文摘要 · Abstract (English)

Magnetic resonance imaging (MRI) quality assessment is crucial for clinical decision-making, yet remains challenging due to data scarcity and protocol variability. Traditional approaches face fundamental trade-offs: signal-based methods like MRIQC provide quantitative metrics but lack semantic understanding, while deep learning approaches achieve high accuracy but sacrifice interpretability. To address these limitations, we introduce the Multimodal MRI Quality Assessment (MMRQA) framework, pioneering the integration of multimodal large language models (MLLMs) with acquisition-aware signal processing. MMRQA combines three key innovations: robust metric extraction via MRQy augmented with simulated artifacts, structured transformation of metrics into question-answer pairs using Qwen, and parameter-efficient fusion through Low-Rank Adaptation (LoRA) of LLaVA-OneVision. Evaluated on MR-ART, FastMRI, and MyConnectome benchmarks, MMRQA achieves state-of-the-art performance with strong zero-shot generalization, as validated by comprehensive ablation studies. By bridging quantitative analysis with semantic reasoning, our framework generates clinically interpretable outputs that enhance quality control in dynamic medical settings.

医学影像多模态大模型质量评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。