arXiv:2510.03232cs.CV2025-10

用少量标注数据让多模态大模型学会医学影像等专业领域问答。

LEAML: Label-Efficient Adaptation to Out-of-Distribution Visual Tasks for Multimodal Large Language Models

  • 用无标签图像生成伪问答对,结合图文描述蒸馏提升效率。
  • 仅更新关键神经元,实现低资源下精准知识迁移。
  • 适合医疗、体育等标注稀缺的垂直领域应用。

多模态大语言模型在通用视觉基准上表现优异,但在医疗影像等专业领域的分布外(OOD)任务中因标注数据稀少而表现不佳。本文提出LEAML框架,利用有限标注的VQA样本与大量无标签图像进行标签高效适配。通过问答生成器生成领域相关的伪问答对,并以图文描述蒸馏进行正则化。关键创新在于仅选择性更新与问答最相关的神经元,使生成器在蒸馏过程中高效获取领域知识。在胃肠道内镜和体育VQA任务上的实验表明,相比标准微调,LEAML在极低监督条件下持续表现更优,验证了该框架的有效性。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have achieved strong performance on general visual benchmarks but struggle with out-of-distribution (OOD) tasks in specialized domains such as medical imaging, where labeled data is limited and expensive. We introduce LEAML, a label-efficient adaptation framework that leverages both scarce labeled VQA samples and abundant unlabeled images. Our approach generates domain-relevant pseudo question-answer pairs for unlabeled data using a QA generator regularized by caption distillation. Importantly, we selectively update only those neurons most relevant to question-answering, enabling the QA Generator to efficiently acquire domain-specific knowledge during distillation. Experiments on gastrointestinal endoscopy and sports VQA demonstrate that LEAML consistently outperforms standard fine-tuning under minimal supervision, highlighting the effectiveness of our proposed LEAML framework.

多模态低资源医学影像问答生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。