视觉语言模型会无意识信任新闻来源的品牌,哪怕内容相反。
Brand-as-Memory: Vision-Language Models Encode Causal, Mechanistically Localizable Credibility Priors for News Sources

- 用源标识符测试模型对新闻可信度的判断,发现品牌名、标志或域名可直接触发信任
- 更换文章标题上的品牌信息,可信度评分变化达11个对数几率单位,与专业评级高度相关
- 可定位到模型第19-21层的特定神经元,且该机制在多个模型中重复出现
视觉语言模型(VLMs)越来越多地通过图像阅读新闻和网页内容,其中出版方身份以视觉形式呈现。我们发现,这些模型对新闻来源存在强烈的可信度先验,且基于出版方身份。研究从三个维度展开:(i) 跨模型基准测试。提出CueTrust诊断工具,通过源覆盖指数(SOI)衡量表面线索如何压倒内容证据。在七种VLMs和五类线索中,脆弱性特征依赖于模型与规模,且覆盖效应具有出口身份特异性、编码不变性——仅由刊头名称、标志图像或域名触发,而作者姓名、文中权威表述或页面布局则为有效负控。 (ii) 机制解释。针对品牌线索,给出完整机制分析:仅替换刊头即可使可信度跨越约11个对数几率单位,与媒体偏见/事实核查(Media Bias/Fact Check)评分高度相关(ρ = 0.88)。该先验为双编码(名称与图像),随模型规模增强,因果形成于第19-21层,由可解释的稀疏自编码器特征携带,并在另一模型族中复现于相同相对位置。其覆盖内容的效应约为1.8倍信号强度,存在于共享路径而非特权通路;定向调控该局部方向可降低覆盖效应41%,且泛化至未见过的出版方,证实该先验为因果使用,非单纯可解码。部署中的VLM可能因此忽视眼前证据而盲目信赖来源身份,这种可靠性缺陷可在多模型中测量、定位并因果探测。研究已发布刺激数据集与CueTrust工具。
原文摘要 · Abstract (English)
Vision-language models (VLMs) increasingly read news and web content as images, where the publisher's identity is visually present. We show that VLMs carry a strong source-credibility prior keyed on outlet identity, and study it along three axes. (i) Cross-model benchmark. We introduce CueTrust, a cross-model diagnostic that measures which surface source cue overrides an article's content evidence via a Source-Override Index (SOI). Across seven VLMs and five cues, the vulnerability profile is model- and scale-dependent, and the override is outlet-identity-specific and encoding-invariant, firing from the masthead name, the logo image, or the bare domain, but not from a named author, in-text authority, or page layout (clean negative controls). (ii) Mechanistic account. For the brand cue, we give a full mechanistic account: swapping only the masthead moves credibility across an approximately 11 log-odds range that tracks professional ratings (rho = 0.88 with Media Bias/Fact Check). The prior is dual-coded (name and logo), strengthens with scale, is causally formed at layers 19-21, carried by interpretable seed-stable sparse-autoencoder features, and recurs at the same relative locus in a second model family. It overrides content (about 1.8x) as a signal-magnitude effect within a shared pathway, not a privileged route. Steering the localized direction selectively reduces the override (41% reduction) and generalizes to held-out outlets, confirming the prior is causally used, not merely decodable. Deployed VLMs may thus defer to source identity over the evidence in front of them, a reliability failure we can measure across models, localize, and causally probe. We release the stimulus suite and CueTrust.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。