arXiv:2501.03012cs.AIcs.CL2025-01ICCV被引 16

通过概念映射分析多模态大模型微调中的表征变化,实现行为可解释与可控调节。

Analyzing Finetuning Representation Shift for Multimodal LLMs Steering

  • 将隐藏状态映射到可解释的视觉/文本概念,追踪微调前后语义动态变化。
  • 提出移位向量可低成本恢复微调后概念,且能有效捕捉概念偏差。
  • 无需重新训练,直接用于模型去偏与安全控制,适合开发者和研究者使用。

多模态大语言模型(MLLMs)在理解多模态输入方面已达到高水平性能。然而,理解和解释这些复杂模型的行为极具挑战性,尤其在微调过程中可能发生的动态表征变化,或因数据集间协变量偏移引发的问题。本文采用概念级分析方法来理解MLLM。具体而言,我们提出将隐藏状态映射到可解释的视觉与文本概念,从而更高效地比较原始模型与微调后模型之间的语义动态,揭示微调过程中的概念改变及潜在偏见。我们还展示了利用移位向量捕捉概念变化的方法。这些移位向量使我们能够在原始模型中通过简单的、计算成本低的加性概念移位,恢复微调后的概念。最终,我们的发现可直接应用于MLLM的可控性设计,可用于模型去偏以及强化输出安全性。总体而言,我们提出了一种全新的、无需训练、即插即用的MLLM行为可解释与控制框架,代码已公开。

原文摘要 · Abstract (English)

Multimodal LLMs (MLLMs) have reached remarkable levels of proficiency in understanding multimodal inputs. However, understanding and interpreting the behavior of such complex models is a challenging task, not to mention the dynamic shifts that may occur during fine-tuning, or due to covariate shift between datasets. In this work, we apply concept-level analysis towards MLLM understanding. More specifically, we propose to map hidden states to interpretable visual and textual concepts. This enables us to more efficiently compare certain semantic dynamics, such as the shift from an original and fine-tuned model, revealing concept alteration and potential biases that may occur during fine-tuning. We also demonstrate the use of shift vectors to capture these concepts changes. These shift vectors allow us to recover fine-tuned concepts by applying simple, computationally inexpensive additive concept shifts in the original model. Finally, our findings also have direct applications for MLLM steering, which can be used for model debiasing as well as enforcing safety in MLLM output. All in all, we propose a novel, training-free, ready-to-use framework for MLLM behavior interpretability and control. Our implementation is publicly available.

多模态可解释性微调模型控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。