视觉指令微调将图像特征嵌入语言模型的语义中间层,实现跨模态对齐。
Visual Instruction Tuning Aligns Modalities through Abstraction

- 通过指令微调将图像特征直接注入语言模型的语义中间层
- 中间层是多模态任务的核心,对各类基准测试性能至关重要
- 仅微调中间层即可保持性能,显著降低训练时间
视觉指令微调能有效将预训练大语言模型(LLM)适配为处理图像与文本信息。然而,视觉特征如何融入LLM分层抽象结构尚不明确。我们针对多种视觉-语言架构发现,指令微调主要作为桥梁,将视觉特征直接嵌入到语言模型的中间语义层,跳过专注于单模态处理的早期层。通过探测分析和因果干预,我们证实这些中间层是多模态处理的语义核心,在广泛多模态基准上起关键作用。此外,通过对比语义等价的视觉与文本表示几何结构,发现微调扩展并强化了现有抽象阶段,使视觉特征与已有文本特征对齐。最后,我们通过仅在中间层进行微调验证其功能:该策略在视觉主导基准上保持全量微调性能,同时显著减少训练时间。结果表明,多模态融合是一种由语言模型内部抽象机制重构驱动的局部现象。
原文摘要 · Abstract (English)
Visual instruction tuning effectively adapts a pre-trained Large Language Model (LLM) to process image information alongside text. Yet, it remains unclear how visual features are embedded into the layer-wise hierarchy of abstractions of the LLM backbone. Across a diverse set of vision-language architectures, we show that instruction tuning primarily serves as a bridge, embedding visual features directly into the intermediate semantic layers of the LLM, bypassing the early layers devoted to unimodal processing. With probing analyses and causal interventions, we show that these intermediate layers are the semantic core of vision-language processing and play a critical role in the performance on a broad set of multimodal benchmarks. In addition, by comparing the geometry of semantically equivalent visual and textual representations, we find that fine-tuning extends and strengthens the existing abstraction phase, aligning visual features with pre-existing textual ones. Finally, we confirm the functional role of this localized alignment by restricting fine-tuning to intermediate layers alone: this strategy preserves the performance of full fine-tuning on vision-centric benchmarks while reducing training time. Our results suggest that multimodal integration is a localized phenomenon driven by the repurposing of the internal abstraction engine of the LLM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。