通过可控生成与吉布斯采样,直接探测多模态大模型的感知先验。
Probing Perceptual Priors of MLLMs via Gibbs Sampling with Interpretable Generative Controls

- 用生成模型沿可控制轴生成图像,以吉布斯采样探索模型的内在偏好分布。
- 在人脸可信度、艺术品廉价感等任务中发现典型偏见与新奇先验。
- 适合研究模型偏差、视觉语言模型行为机制的学者使用。
模型在任务中的表现由输入和其自带的先验共同决定,即模型隐含期待的刺激分布。传统可解释性研究通常固定输入,从内部结构或输出变化分析模型,但无法重建先验分布——内部结构仅反映模型能表示什么,而非它期望什么;而固定输入集则遗漏了大部分可能输入空间。尤其在真实场景中(如视觉语言模型看到的图像),输入空间维度极高且多样,这些先验仍是影响实际行为却未被充分理解的关键成分。本文提出一种直接采样模型感知先验分布的方法:通过生成模型沿可控轴生成刺激,并以待研究模型作为判别器运行吉布斯采样。我们将其应用于多个类别与目标变量(如人脸可信度、艺术作品廉价感),成功识别出既有的常见偏见以及通过直接提示无法发现的新型先验,值得进一步探究其下游影响。
原文摘要 · Abstract (English)
A model's behavior on a task is jointly determined by the input it receives and the prior it brings in, i.e. the distribution over stimuli it implicitly expects. Interpretability research has traditionally studied models by holding inputs fixed and examining model responses either mechanistically, probing how internal structure represents inputs, or behaviorally, measuring how variation in inputs leads to variation in outputs. Neither reconstructs the prior distribution itself, since internal structure shows what a model can represent, not what it expects, and any fixed stimulus set leaves most of the possible input space unseen. In particular, such an input space in real-world settings, such as images seen by VLMs, is extremely high-dimensional and diverse. These priors thus remain a poorly understood component of models that nonetheless influence real-world behavior. We propose a method to sample from models' perceptual prior distributions directly, by steering a generative model to produce stimuli along controllable axes and running Gibbs sampling over that space with the model under study as the judge. We apply this to a variety of categories and target variables (such as trustworthiness in faces and cheapness in art images) and recover both canonical biases and surprising novel priors invisible to direct prompting, warranting further investigation of their downstream effects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。