arXiv:2603.19482cs.CV2026-03

无需人工指令,用图像描述对训练医学视觉语言模型

Instruction-Free Tuning of Large Vision Language Models for Medical Instruction Following

  • 用图像描述对替代人工指令进行微调
  • 在多个医学数据集上达到顶尖准确率
  • 适合缺乏专业标注资源的医疗AI场景

大型视觉语言模型(LVLM)在多种任务中表现优异,主要得益于视觉指令微调——即使用精心构建的图像-指令-输出三元组数据集进行微调。然而,在医学领域,由于需要专业专家知识,构建大规模高质量指令数据集尤为困难。为此,我们提出一种无指令微调方法,仅依赖图像-描述对进行训练。具体而言,引入动量代理指令替代人工文本指令,既保持预训练模型的指令遵循能力,又促进推理时仍有效的参数更新。因此,微调后的模型能灵活响应特定领域指令,即使训练时未显式提供指令。此外,采用响应打乱策略缓解模型对前序词的过度依赖,提升微调效果。该方法在SKINCON、WBCAtt、CBIS和MIMIC-CXR等多个多选题视觉问答任务上达到当前最优性能,显著提升了医学领域LVLM的微调效率。

原文摘要 · Abstract (English)

Large vision language models (LVLMs) have demonstrated impressive performance across a wide range of tasks. These capabilities largely stem from visual instruction tuning, which fine-tunes models on datasets consisting of curated image-instruction-output triplets. However, in the medical domain, constructing large-scale, high-quality instruction datasets is particularly challenging due to the need for specialized expert knowledge. To address this issue, we propose an instruction-free tuning approach that reduces reliance on handcrafted instructions, leveraging only image-description pairs for fine-tuning. Specifically, we introduce a momentum proxy instruction as a replacement for curated text instructions, which preserves the instruction-following capability of the pre-trained LVLM while promoting updates to parameters that remain valid during inference. Consequently, the fine-tuned LVLM can flexibly respond to domain-specific instructions, even though explicit instructions are absent during fine-tuning. Additionally, we incorporate a response shuffling strategy to mitigate the model's over-reliance on previous words, facilitating more effective fine-tuning. Our approach achieves state-of-the-art accuracy on multiple-choice visual question answering tasks across SKINCON, WBCAtt, CBIS, and MIMIC-CXR datasets, significantly enhancing the fine-tuning efficiency of LVLMs in medical domains.

医学视觉语言模型无指令微调图像描述对数据效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。