发现视觉语言模型输出稳定但内部表示不稳,可能误导评估结果。
Same Answer, Different Representations: Hidden instability in VLMs
- 提出新评估框架,检测模型内部表征漂移与频域敏感性
- 大模型准确率高但更易受干扰,决策边界更锋利脆弱
- 相同答案下内部变化巨大,适合关注模型可靠性研究者
视觉语言模型(VLM)的鲁棒性通常通过输出不变性来评估,隐含假设是稳定预测反映稳定的多模态处理。本文指出该假设不足,提出一个兼顾表征感知与频率感知的评估框架,测量内部嵌入漂移、频谱敏感性及结构平滑性(视觉标记的空间一致性),并结合标准标签指标。在SEEDBench、MMMU和POPE数据集上对现代VLMs的测试揭示三种不同失效模式:第一,模型常在输出不变时发生显著内部表征漂移,例如文本叠加扰动下,漂移幅度接近跨图像差异,表明表征移动至通常由无关输入占据的区域;第二,鲁棒性不随规模提升,更大模型虽准确率更高,但敏感度相当或更高,符合更尖锐却更脆弱的决策边界;第三,扰动影响任务不同:破坏粗细视觉线索融合会损害推理,但在幻觉基准上反而降低假阳性,使模型生成更保守答案。
原文摘要 · Abstract (English)
The robustness of Vision Language Models (VLMs) is commonly assessed through output-level invariance, implicitly assuming that stable predictions reflect stable multimodal processing. In this work, we argue that this assumption is insufficient. We introduce a representation-aware and frequency-aware evaluation framework that measures internal embedding drift, spectral sensitivity, and structural smoothness (spatial consistency of vision tokens), alongside standard label-based metrics. Applying this framework to modern VLMs across the SEEDBench, MMMU, and POPE datasets reveals three distinct failure modes. First, models frequently preserve predicted answers while undergoing substantial internal representation drift; for perturbations such as text overlays, this drift approaches the magnitude of inter-image variability, indicating that representations move to regions typically occupied by unrelated inputs despite unchanged outputs. Second, robustness does not improve with scale; larger models achieve higher accuracy but exhibit equal or greater sensitivity, consistent with sharper yet more fragile decision boundaries. Third, we find that perturbations affect tasks differently: they harm reasoning when they disrupt how models combine coarse and fine visual cues, but on the hallucination benchmarks, they can reduce false positives by making models generate more conservative answers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。