arXiv:2504.09967cs.CVcs.AI2025-04被引 1

构建图像为中心的多标注数据集,提升医学大模型多任务能力。

Enhancing Multi-task Learning Capability of Medical Generalist Foundation Model via Image-centric Multi-annotation Data

  • 以图像为中心构建多标注数据,每张影像平均关联4.1个任务
  • 在7个任务上使模型平均性能提升3.2%至21.05%
  • 适合追求临床实用性的医学多模态模型研究者

医学通用基础模型的兴起改变了传统任务专用模型的开发范式,旨在通过大规模医学数据联合训练更好地处理多种任务。然而,现有进展多聚焦于数据量扩展或架构优化,忽视了从数据层面重思多任务学习。单纯整合已有数据资源会导致图像与任务的分散对齐,难以培养全面的图像理解能力,也偏离临床所需的多维度影像解读需求。本文提出首个从数据构建层面增强医学多模态大语言模型(MLLMs)多任务学习能力的图像为中心多标注X光数据集(IMAX)。IMAX具备以下特点:1)高质量数据筛选,涵盖超过35.4万条适用于七项不同医疗任务的数据;2)以图像为中心的密集标注,每张X光片平均关联4.10个任务和7.46条训练样本,确保每张图像的多任务表征丰富性。相较于通用的去中心化多标注X光数据集(DMAX),IMAX在七个开源先进医学MLLMs上均实现3.20%至21.05%的显著多任务平均性能提升。此外,我们分析了IMAX与DMAX训练过程中的统计模式差异,探索优化动态与多任务性能间的潜在关联。最后,基于IMAX的数据构建核心思想,提出一种优化的基于DMAX的训练策略,以缓解实际场景中获取高质量IMAX数据的困境。

原文摘要 · Abstract (English)

The emergence of medical generalist foundation models has revolutionized conventional task-specific model development paradigms, aiming to better handle multiple tasks through joint training on large-scale medical datasets. However, recent advances prioritize simple data scaling or architectural component enhancement, while neglecting to re-examine multi-task learning from a data-centric perspective. Critically, simply aggregating existing data resources leads to decentralized image-task alignment, which fails to cultivate comprehensive image understanding or align with clinical needs for multi-dimensional image interpretation. In this paper, we introduce the image-centric multi-annotation X-ray dataset (IMAX), the first attempt to enhance the multi-task learning capabilities of medical multi-modal large language models (MLLMs) from the data construction level. To be specific, IMAX is featured from the following attributes: 1) High-quality data curation. A comprehensive collection of more than 354K entries applicable to seven different medical tasks. 2) Image-centric dense annotation. Each X-ray image is associated with an average of 4.10 tasks and 7.46 training entries, ensuring multi-task representation richness per image. Compared to the general decentralized multi-annotation X-ray dataset (DMAX), IMAX consistently demonstrates significant multi-task average performance gains ranging from 3.20% to 21.05% across seven open-source state-of-the-art medical MLLMs. Moreover, we investigate differences in statistical patterns exhibited by IMAX and DMAX training processes, exploring potential correlations between optimization dynamics and multi-task performance. Finally, leveraging the core concept of IMAX data construction, we propose an optimized DMAX-based training strategy to alleviate the dilemma of obtaining high-quality IMAX data in practical scenarios.

医学多模态多任务学习数据构建图像标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。