提出无需训练的校准框架,提升医疗多模态大模型对真实噪声的鲁棒性。
Perceive and Calibrate: Analyzing and Enhancing Robustness of Medical Multi-Modal Large Language Models
- 基于感知-校准原则,利用模型自身能力进行跨模态去噪
- 在11种噪声下实现多模态性能领先,优于现有方法
- 适合医疗AI安全落地,尤其关注临床真实场景的模型可靠性
医学多模态大语言模型(MLLMs)展现出良好的临床表现,但其对成像伪影、文本错误等现实输入扰动敏感,严重制约临床应用。现有研究缺乏对医学领域多模态噪声影响的系统分析,且多数工作仅关注文本模态,依赖昂贵微调,难以应对医学复杂噪声与严格安全标准。为此,本文系统分析了视觉与文本模态中多种扰动对医疗MLLMs的影响。基于发现,提出无需训练的内在增强多模态校准(IMC)框架,遵循感知-校准原则,提升跨模态鲁棒性。针对图像,设计扰动感知去噪校准(PDC),利用模型自有的视觉编码器识别噪声并进行原型引导特征校准;针对文本,构建自生成多智能体系统(SMS),通过智能体协作利用模型自评估能力修复噪声文本。构建包含11类噪声的基准,在2个数据集上验证,实验表明该方法在多模态上达到最新性能,具备提升实际临床场景中模型鲁棒性的潜力。
原文摘要 · Abstract (English)
Medical Multi-modal Large Language Models (MLLMs) have shown promising clinical performance. However, their sensitivity to real-world input perturbations, such as imaging artifacts and textual errors, critically undermines their clinical applicability. Systematic analysis of such noise impact on medical MLLMs remains largely unexplored. Furthermore, while several works have investigated the MLLMs' robustness in general domains, they primarily focus on text modality and rely on costly fine-tuning. They are inadequate to address the complex noise patterns and fulfill the strict safety standards in medicine. To bridge this gap, this work systematically analyzes the impact of various perturbations on medical MLLMs across both visual and textual modalities. Building on our findings, we introduce a training-free Inherent-enhanced Multi-modal Calibration (IMC) framework that leverages MLLMs' inherent denoising capabilities following the perceive-and-calibrate principle for cross-modal robustness enhancement. For the visual modality, we propose a Perturbation-aware Denoising Calibration (PDC) which leverages MLLMs' own vision encoder to identify noise patterns and perform prototype-guided feature calibration. For text denoising, we design a Self-instantiated Multi-agent System (SMS) that exploits the MLLMs' self-assessment capabilities to refine noisy text through a cooperative hierarchy of agents. We construct a benchmark containing 11 types of noise across both image and text modalities on 2 datasets. Experimental results demonstrate our method achieves the state-of-the-art performance across multiple modalities, showing potential to enhance MLLMs' robustness in real clinical scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。