arXiv:2503.00723cs.LG2025-03ICLR被引 21

通过直接编辑多模态表示,实现高效且可解释的模型控制。

Re-Imagining Multimodal Instruction Tuning: A Representation View

  • 不微调模型参数,而是直接修改语义丰富的多模态表示。
  • 在MME上取得1580.40分,仅需0.03%参数量。
  • 支持对模型行为进行直观可控的编辑,适合需要可解释性的场景。

多模态指令微调通过在预训练大型多模态模型(LMMs)上使用指令跟随数据进行微调,已证明能有效实现零样本泛化。然而,随着LMM规模持续扩大,全量微调变得高度参数密集。尽管已有参数高效微调(PEFT)方法降低可调参数数量,但与全量微调相比仍存在显著性能差距。此外,现有PEFT方法往往参数量大,难以解释和控制。为此,我们提出多模态表示微调(MRT),聚焦于直接编辑语义丰富的多模态表示,以实现强性能并提供对LMM的直观控制。实验表明,该方法在多个基准上超越当前最先进基线,性能提升显著(如在MME上达1580.40分),同时仅需极少可调参数(如0.03%)。此外,我们在多模态表示中对特定标记进行编辑实验,证明直接操纵这些表示可实现简单而有效的网络行为控制。

原文摘要 · Abstract (English)

Multimodal instruction tuning has proven to be an effective strategy for achieving zero-shot generalization by fine-tuning pre-trained Large Multimodal Models (LMMs) with instruction-following data. However, as the scale of LMMs continues to grow, fully fine-tuning these models has become highly parameter-intensive. Although Parameter-Efficient Fine-Tuning (PEFT) methods have been introduced to reduce the number of tunable parameters, a significant performance gap remains compared to full fine-tuning. Furthermore, existing PEFT approaches are often highly parameterized, making them difficult to interpret and control. In light of this, we introduce Multimodal Representation Tuning (MRT), a novel approach that focuses on directly editing semantically rich multimodal representations to achieve strong performance and provide intuitive control over LMMs. Empirical results show that our method surpasses current state-of-the-art baselines with significant performance gains (e.g., 1580.40 MME score) while requiring substantially fewer tunable parameters (e.g., 0.03% parameters). Additionally, we conduct experiments on editing instrumental tokens within multimodal representations, demonstrating that direct manipulation of these representations enables simple yet effective control over network behavior.

多模态参数高效表示编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。