揭示视觉语言模型中表示工程为何有效,提升AI透明与可控性。
Why Representation Engineering Works: A Theoretical and Empirical Study in Vision-Language Models
- 用主特征向量解释多模态表示的层间稳定性机制。
- 实证验证了表示工程在视觉语言模型中的广泛适用性。
- 为增强AI鲁棒性、公平性和透明性提供理论框架。
表示工程(RepE)作为一种提升AI透明度的强大范式,聚焦于高层表示而非单个神经元或电路,已在大语言模型中证明能有效增强可解释性与控制力。然而,在视觉语言模型(VLMs)中,视觉输入常压倒事实性语言知识,导致与现实矛盾的幻觉输出。为此,本文首次将RepE扩展至VLMs,分析多模态表示的保持与演化过程。基于发现并借鉴成功案例,构建理论框架,利用主特征向量解释神经活动跨层稳定性,揭示RepE的内在机制。通过实证验证这些内在属性,展示其普遍适用性与重要性。本工作将RepE从描述性工具转变为结构化理论框架,为提升AI鲁棒性、公平性与透明性开辟新方向。
原文摘要 · Abstract (English)
Representation Engineering (RepE) has emerged as a powerful paradigm for enhancing AI transparency by focusing on high-level representations rather than individual neurons or circuits. It has proven effective in improving interpretability and control, showing that representations can emerge, propagate, and shape final model outputs in large language models (LLMs). However, in Vision-Language Models (VLMs), visual input can override factual linguistic knowledge, leading to hallucinated responses that contradict reality. To address this challenge, we make the first attempt to extend RepE to VLMs, analyzing how multimodal representations are preserved and transformed. Building on our findings and drawing inspiration from successful RepE applications, we develop a theoretical framework that explains the stability of neural activity across layers using the principal eigenvector, uncovering the underlying mechanism of RepE. We empirically validate these instrinsic properties, demonstrating their broad applicability and significance. By bridging theoretical insights with empirical validation, this work transforms RepE from a descriptive tool into a structured theoretical framework, opening new directions for improving AI robustness, fairness, and transparency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。