高分辨率图像对生物医学多模态大模型性能至关重要
The Impact of Image Resolution on Biomedical Multimodal Large Language Models
- 用原始分辨率训练和推理可显著提升模型表现
- 训练与推理分辨率不一致会严重降低性能
- 混合分辨率训练能平衡计算成本与效果
成像技术是生物医学研究和现代医学的基础,需分析跨模态的高分辨率图像。尽管多模态大语言模型(MLLM)在生物医学图像分析中展现出潜力,但多数模型基于通用数据集中的低分辨率图像设计,可能导致关键信息丢失。本文研究图像分辨率对生物医学领域MLLM性能的影响,发现:(1) 原始分辨率的训练与推理显著提升多个任务的表现;(2) 训练与推理分辨率不匹配会严重损害性能;(3) 混合分辨率训练能有效缓解分辨率错配问题,并在计算约束与性能需求间取得平衡。基于这些发现,建议优先采用原生分辨率推理和混合分辨率数据集,以优化生物医学MLLM在科研与临床应用中的转化潜力。
原文摘要 · Abstract (English)
Imaging technologies are fundamental to biomedical research and modern medicine, requiring analysis of high-resolution images across various modalities. While multimodal large language models (MLLMs) show promise for biomedical image analysis, most are designed for low-resolution images from general-purpose datasets, risking critical information loss. We investigate how image resolution affects MLLM performance in biomedical applications and demonstrate that: (1) native-resolution training and inference significantly improve performance across multiple tasks, (2) misalignment between training and inference resolutions severely degrades performance, and (3) mixed-resolution training effectively mitigates misalignment and balances computational constraints with performance requirements. Based on these findings, we recommend prioritizing native-resolution inference and mixed-resolution datasets to optimize biomedical MLLMs for transformative impact in scientific research and clinical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。