arXiv:2502.11427cs.CLcs.CV2025-02

无需图像指令,仅用文本和描述就能训练视觉语言模型。

Do we Really Need Visual Instructions? Towards Visual Instruction-Free Fine-tuning for Large Vision-Language Models

  • 用纯文本指令和图像描述训练,分离学习推理与视觉能力
  • 在多个视觉推理任务上达到顶尖性能,且训练数据更少
  • 适合资源有限或想减少标注成本的研究者

视觉指令微调已成为激发大型视觉语言模型(LVLM)多模态任务解决能力的主流技术。尽管效果显著,但视觉指令需以图像为输入,导致难以继承基础大语言模型的任务求解能力,且大规模数据收集成本高昂。为此,我们提出ViFT——一种无需视觉指令的LVLM微调框架。在训练中,仅需文本指令和图像描述数据,分别学习任务求解与视觉感知能力。推理时,融合文本与图像输入的表示,实现多模态任务完成。实验表明,ViFT在多个视觉推理与视觉指令遵循基准上达到当前最优表现,且所需训练数据更少。代码与数据将公开发布。

原文摘要 · Abstract (English)

Visual instruction tuning has become the predominant technology in eliciting the multimodal task-solving capabilities of large vision-language models (LVLMs). Despite the success, as visual instructions require images as the input, it would leave the gap in inheriting the task-solving capabilities from the backbone LLMs, and make it costly to collect a large-scale dataset. To address it, we propose ViFT, a visual instruction-free fine-tuning framework for LVLMs. In ViFT, we only require the text-only instructions and image caption data during training, to separately learn the task-solving and visual perception abilities. During inference, we extract and combine the representations of the text and image inputs, for fusing the two abilities to fulfill multimodal tasks. Experimental results demonstrate that ViFT can achieve state-of-the-art performance on several visual reasoning and visual instruction following benchmarks, with rather less training data. Our code and data will be publicly released.

视觉语言模型无图像指令微调框架低资源训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。