arXiv:2601.02443cs.CVcs.AI2026-01被引 1

医学影像诊断中,大模型不如视觉编码器直接分类准确。

Evaluating the Diagnostic Classification Ability of Multimodal Large Language Models: Insights from the Osteoarthritis Initiative

  • 用多模态大模型处理膝骨关节炎影像分类,发现仅靠视觉编码器效果更好。
  • 小而均衡的数据集(500张)微调比大而不均衡数据(5778张)更有效。
  • 大模型更适合生成解释报告,而非直接做高精度诊断分类。

多模态大语言模型(MLLM)在医学视觉问答和报告生成上表现优异,但其生成与解释能力难以可靠迁移至疾病特异性分类任务。本文评估了多种MLLM架构在膝骨关节炎(OA)X光片分类中的表现,该任务在现有医学MLLM基准中仍被严重低估。通过系统性消融实验,考察了视觉编码器、连接模块和大语言模型(LLM)在不同训练策略下的贡献。结果表明,在分类任务中,仅使用训练后的视觉编码器即可超越完整MLLM流程的准确率;微调LLM并未带来显著提升,而提示引导已足够。此外,在小规模且类别均衡的数据集(500张图像)上进行LoRA微调的效果优于在大规模但类别不平衡的数据集(5778张图像)上的训练,说明数据平衡与质量对本任务更为关键。研究建议:在特定医学分类任务中,应优先优化视觉编码器并精心设计数据集,而非依赖大模型作为主要分类器。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) show promising performance on medical visual question answering (VQA) and report generation, but these generation and explanation abilities do not reliably transfer to disease-specific classification. We evaluated MLLM architectures on knee osteoarthritis (OA) radiograph classification, which remains underrepresented in existing medical MLLM benchmarks, even though knee OA affects an estimated 300 to 400 million people worldwide. Through systematic ablation studies manipulating the vision encoder, the connector, and the large language model (LLM) across diverse training strategies, we measured each component's contribution to diagnostic accuracy. In our classification task, a trained vision encoder alone could outperform full MLLM pipelines in classification accuracy and fine-tuning the LLM provided no meaningful improvement over prompt-based guidance. And LoRA fine-tuning on a small, class-balanced dataset (500 images) gave better results than training on a much larger but class-imbalanced set (5,778 images), indicating that data balance and quality can matter more than raw scale for this task. These findings suggest that for domain-specific medical classification, LLMs are more effective as interpreters and report generators rather than as primary classifiers. Therefore, the MLLM architecture appears less suitable for medical image diagnostic classification tasks that demand high certainty. We recommend prioritizing vision encoder optimization and careful dataset curation when developing clinically applicable systems.

医学影像多模态分类数据质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。