arXiv:2510.18304cs.CVcs.CL2025-10被引 1

高分辨率图像对生物医学多模态大模型性能至关重要

The Impact of Image Resolution on Biomedical Multimodal Large Language Models

  • 用原始分辨率训练和推理可显著提升模型表现
  • 训练与推理分辨率不一致会严重降低性能
  • 混合分辨率训练能平衡计算成本与效果

成像技术是生物医学研究和现代医学的基础,需分析跨模态的高分辨率图像。尽管多模态大语言模型(MLLM)在生物医学图像分析中展现出潜力,但多数模型基于通用数据集中的低分辨率图像设计,可能导致关键信息丢失。本文研究图像分辨率对生物医学领域MLLM性能的影响,发现:(1) 原始分辨率的训练与推理显著提升多个任务的表现;(2) 训练与推理分辨率不匹配会严重损害性能;(3) 混合分辨率训练能有效缓解分辨率错配问题,并在计算约束与性能需求间取得平衡。基于这些发现,建议优先采用原生分辨率推理和混合分辨率数据集,以优化生物医学MLLM在科研与临床应用中的转化潜力。

原文摘要 · Abstract (English)

Imaging technologies are fundamental to biomedical research and modern medicine, requiring analysis of high-resolution images across various modalities. While multimodal large language models (MLLMs) show promise for biomedical image analysis, most are designed for low-resolution images from general-purpose datasets, risking critical information loss. We investigate how image resolution affects MLLM performance in biomedical applications and demonstrate that: (1) native-resolution training and inference significantly improve performance across multiple tasks, (2) misalignment between training and inference resolutions severely degrades performance, and (3) mixed-resolution training effectively mitigates misalignment and balances computational constraints with performance requirements. Based on these findings, we recommend prioritizing native-resolution inference and mixed-resolution datasets to optimize biomedical MLLMs for transformative impact in scientific research and clinical applications.

多模态模型图像分辨率生物医学AIMLLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。