提出多模态模型中双通道语境牵连的测评工具,揭示视觉与语言如何独立误导模型输出。
ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models
- 设计双流评测框架,分别测试文本和图像对模型输出的干扰
- 构建1500条标注数据,覆盖8类情境与真假关系的复杂组合
- 提供可复用的评估协议,助力研究多模态模型可靠性
上下文牵连指模型在未考虑内容相关性、真实性或意义的情况下,受输入中附加上下文影响而改变输出。这一现象已在单模态语言模型中被发现并机制化。然而,在视觉-语言模型(VLMs)中,其表现形式尚不明确,且缺乏专用评测工具。本文认为,研究VLM中的牵连不能简单将纯文本基准移植至多模态场景,而需一个基于双模态、结构化的评估体系,其条件应围绕具体对象展开——即文本流中的图像内容,以及视觉流中的文本查询。我们提出ENTRAP-VL(视觉语言牵连评估探针),一个手工标注的1500项数据集,涵盖8个类别,按两个轴分类:上下文与目标对象的关系及其与真实性的关联。该数据集分为文本牵连流(8种情境)和视觉牵连流(3种情境)。我们不直接测量特定模型,而是提供工具、分类框架及评估协议,供社区开展严谨研究。数据集与文档将公开发布。
原文摘要 · Abstract (English)
Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that context is relevant, true, or even meaningful. Recently, it has been identified and given a mechanistic account in unimodal language models. Whether and how it manifests in vision-language models (VLMs) is, by contrast, largely unexamined, and the field lacks a purpose-built instrument with which to investigate it. We take the position that studying contextual entrainment in VLMs requires more than porting an existing text-only benchmark to the multimodal setting: it requires a taxonomically structured, dual-modality instrument whose conditions are constructed around the item at hand (the depicted image in the textual stream, the textual query in the visual stream). We argue that the move to VLMs is substantive rather than incremental. It makes entrainment a dual phenomenon, drivable independently by textual and by visual context, and it opens a veracity distinction (context that is false of the depicted scene yet possible in the world) that has no counterpart in the unimodal, world-knowledge-only formulation of prior work. To make this position concrete and actionable, we introduce ENTRAP-VL (ENTRainment Assessment Probe for Vision and Language), a manually curated dataset of 1,500 items across eight categories, organized by a taxonomy that spans two axes, i.e., the association of context with the item and its relationship to truth, and split into a textual-entrainment stream (eight context conditions) and a visual-entrainment stream (three context conditions). We do not claim to measure entrainment in any particular model; we provide the instrument, the taxonomy that motivates it, and the evaluation protocols it enables, so that the community can investigate the phenomenon rigorously. We will release the dataset and its documentation publicly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。