揭示医学多模态大模型图像分类性能下降的四大根源
Lost in the Hype: Revealing and Dissecting the Performance Degradation of Medical Multimodal Large Language Models in Image Classification
- 通过逐层特征探查,定位视觉表征与语义映射中的信息损耗点
- 发现四类失效模式:视觉表征质量差、投影失真、理解不足、语义错配
- 提出量化指标评估特征演化健康度,适用于不同模型和数据集比较
多模态大语言模型(MLLMs)在医学影像分析中广泛应用,但其在图像分类这一基础任务上表现却持续落后于传统深度学习模型,尽管前者拥有更庞大的预训练数据和参数量。本文在三个代表性数据集上对14个开源医学MLLMs进行系统实验,突破表面性能对比,采用特征探查技术逐模块、逐层追踪视觉特征流动,首次明确可视化分类信号在何处被扭曲、稀释或覆盖。研究揭示四类失效模式:1)视觉表征质量受限;2)连接器投影保真度下降;3)大语言模型推理理解不足;4)语义映射错位。同时引入量化评分以衡量特征演化的健康程度,支持跨模型、跨数据集的严谨比较。文章还深入探讨当前医学MLLMs难以实现临床落地的关键障碍,呼吁学界重新审视从期望到实用的漫长路径。
原文摘要 · Abstract (English)
The rise of multimodal large language models (MLLMs) has sparked an unprecedented wave of applications in the field of medical imaging analysis. However, as one of the earliest and most fundamental tasks integrated into this paradigm, medical image classification reveals a sobering reality: state-of-the-art medical MLLMs consistently underperform compared to traditional deep learning models, despite their overwhelming advantages in pre-training data and model parameters. This paradox prompts a critical rethinking: where exactly does the performance degradation originate? In this paper, we conduct extensive experiments on 14 open-source medical MLLMs across three representative image classification datasets. Moving beyond superficial performance benchmarking, we employ feature probing to track the information flow of visual features module-by-module and layer-by-layer throughout the entire MLLM pipeline, enabling explicit visualization of where and how classification signals are distorted, diluted, or overridden. As the first attempt to dissect classification performance degradation in medical MLLMs, our findings reveal four failure modes: 1) quality limitation in visual representation, 2) fidelity loss in connector projection, 3) comprehension deficit in LLM reasoning, and 4) misalignment of semantic mapping. Meanwhile, we introduce quantitative scores that characterize the healthiness of feature evolution, enabling principled comparisons across diverse MLLMs and datasets. Furthermore, we provide insightful discussions centered on the critical barriers that prevent current medical MLLMs from fulfilling their promised clinical potential. We hope that our work provokes rethinking within the community-highlighting that the road from high expectations to clinically deployable MLLMs remains long and winding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。