arXiv:2502.17514cs.LGcs.AI2025-02ICML被引 23

用稀疏自编码器解析多模态模型,提升对齐效果与可解释性。

SAE-V: Interpreting Multimodal Models for Enhanced Alignment

  • 扩展稀疏自编码器至多模态场景,识别跨模态可解释特征
  • 仅用50%数据即可实现超110%性能提升,显著优化对齐效果
  • 内置数据过滤机制,无需额外模型,适合调试和优化多模态模型

随着图像模态的引入,多模态大语言模型(MLLMs)的语义空间比纯文本模型更复杂,导致可解释性更难,对齐也更不稳定,尤其易受低质量数据影响,引发模态不一致、幻觉和偏见。因此,发展MLLM的可解释方法对提升对齐质量至关重要。在纯文本大模型中,稀疏自编码器(SAEs)已展现其解释能力。但将SAEs扩展到多模态场景面临模态融合与跨模态表示分离难题。为此,我们提出SAE-V,一种面向MLLM的机制可解释性框架,通过识别并分析可解释特征及其对应数据,实现对模型行为与数据质量的细粒度解析,深化对跨模态交互与对齐动态的理解。此外,借助跨模态特征加权,SAE-V提供内在数据过滤机制,在不依赖额外模型的前提下增强模型对齐。具体应用于MLLM对齐过程时,基于SAE-V的数据过滤方法可在使用不足50%数据的情况下实现超过110%的性能提升。结果表明,SAE-V能有效增强MLLM的可解释性与对齐能力,揭示其内部运作机制。

原文摘要 · Abstract (English)

With the integration of image modality, the semantic space of multimodal large language models (MLLMs) is more complex than text-only models, making their interpretability more challenging and their alignment less stable, particularly susceptible to low-quality data, which can lead to inconsistencies between modalities, hallucinations, and biased outputs. As a result, developing interpretability methods for MLLMs is crucial for improving alignment quality and efficiency. In text-only LLMs, Sparse Autoencoders (SAEs) have gained attention for their ability to interpret latent representations. However, extending SAEs to multimodal settings presents new challenges due to modality fusion and the difficulty of isolating cross-modal representations. To address these challenges, we introduce SAE-V, a mechanistic interpretability framework that extends the SAE paradigm to MLLMs. By identifying and analyzing interpretable features along with their corresponding data, SAE-V enables fine-grained interpretation of both model behavior and data quality, facilitating a deeper understanding of cross-modal interactions and alignment dynamics. Moreover, by utilizing cross-modal feature weighting, SAE-V provides an intrinsic data filtering mechanism to enhance model alignment without requiring additional models. Specifically, when applied to the alignment process of MLLMs, SAE-V-based data filtering methods could achieve more than 110% performance with less than 50% data. Our results highlight SAE-V's ability to enhance interpretability and alignment in MLLMs, providing insights into their internal mechanisms.

多模态可解释性对齐自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。