无需标注数据,让医疗多模态大模型在测试时自我进化。
Med-Evo: Test-time Self-evolution for Medical Multimodal Large Language Models
- 用无标签数据生成伪标签,引导模型自我优化。
- 在SLAKE数据集上提升10.43%准确率,召回率增4.68%。
- 适合医疗领域数据稀缺场景,推动模型持续进化。
医疗多模态大模型(MLLMs)在多种医疗任务中表现卓越。然而,现有后训练方法如监督微调和强化学习高度依赖大量标注数据,忽视了未标注测试数据的潜力。这一局限在医疗领域尤为突出,因数据敏感性和标注复杂性导致大规模标注数据难以获取。此外,利用测试数据面临如何从无标签样本生成可靠监督信号及保持自进化稳定性的问题。为此,我们提出Med-Evo,首个面向医疗MLLM的自演化框架,采用无标签强化学习,在不需额外标注数据的前提下提升模型性能。该框架引入两项创新:1)特征驱动伪标签(FPL),从所有异构候选响应中识别语义中心点以选择伪标签;2)硬-软奖励(HSR),结合精确匹配、词元级评估与语义相似性提供分层奖励。在三个医学VQA基准和两个基础MLLM上的实验表明,本方法显著优于现有最先进方法,在SLAKE数据集上使用Qwen2.5-VL时,准确率提升10.43%,召回率提升4.68%,验证了方法有效性。
原文摘要 · Abstract (English)
Medical Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities across diverse healthcare tasks. However, current post-training strategies, such as supervised fine-tuning and reinforcement learning, heavily depend on substantial annotated data while overlooking the potential of unlabeled test data for model enhancement. This limitation becomes particularly pronounced in medical domains, where acquiring extensive labeled medical data is difficult due to the strict data sensitivity and annotation complexity. Moreover, leveraging test data poses challenges in generating reliable supervision signals from unlabeled samples and maintaining stable self-evolution. To address these limitations, we propose Med-Evo, the first self-evolution framework for medical MLLMs that utilizes label-free reinforcement learning to promote model performance without requiring additional labeled data. Our framework introduces two key innovations: $1)$ Feature-driven Pseudo Labeling (FPL) that identifies semantic centroids from all heterogeneous candidate responses to select pseudo labels in each rollout, and $2)$ Hard-Soft Reward (HSR) that combines exact match with token-level assessment and semantic similarity to provide hierarchical reward. Experiments on three medical VQA benchmarks and two base MLLMs show clear advantages of our approach over SOTA methods, with significant improvements of 10.43\% accuracy and 4.68\% recall on the SLAKE dataset using Qwen2.5-VL, showing the effectiveness of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。