arXiv:2503.04543cs.CLcs.AI2025-03被引 13

针对多模态大模型微调中的遗忘与泛化难题,提出系统性解决方案。

Keeping Yourself is Important in Downstream Tuning Multimodal Large Language Model

  • 按选择性、增量式、重参数三类方法分类微调策略
  • 实证验证不同方法在多个任务上的性能差异与稳定性
  • 适合关注模型持续学习与下游适配的研究者

多模态大语言模型(MLLMs)融合视觉与语言推理,可完成图像描述、视觉问答等复杂任务。然而在特定应用中表现受限。微调时面临两大挑战:任务专家专化(预训练与目标数据分布偏移限制性能)和开放世界稳定(灾难性遗忘导致通用知识丢失)。本文系统梳理近期微调方法,将其分为三类:(I) 选择性微调,(II) 增量式微调,(III) 重参数化微调。同时在主流MLLM架构与多样化下游任务上进行基准测试,建立标准化评估体系与系统性微调原则。最后指出该领域现存开放问题并提出未来方向。为推动进展,作者开源持续更新的资源库:https://github.com/WenkeHuang/Awesome-MLLM-Tuning。

原文摘要 · Abstract (English)

Multi-modal Large Language Models (MLLMs) integrate visual and linguistic reasoning to address complex tasks such as image captioning and visual question answering. While MLLMs demonstrate remarkable versatility, MLLMs appears limited performance on special applications. But tuning MLLMs for downstream tasks encounters two key challenges: Task-Expert Specialization, where distribution shifts between pre-training and target datasets constrain target performance, and Open-World Stabilization, where catastrophic forgetting erases the model general knowledge. In this work, we systematically review recent advancements in MLLM tuning methodologies, classifying them into three paradigms: (I) Selective Tuning, (II) Additive Tuning, and (III) Reparameterization Tuning. Furthermore, we benchmark these tuning strategies across popular MLLM architectures and diverse downstream tasks to establish standardized evaluation analysis and systematic tuning principles. Finally, we highlight several open challenges in this domain and propose future research directions. To facilitate ongoing progress in this rapidly evolving field, we provide a public repository that continuously tracks developments: https://github.com/WenkeHuang/Awesome-MLLM-Tuning.

多模态模型微调持续学习LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。