提出无需标签的视觉语言模型可靠性检测新方法,揭示错误不可见的本质原因。
Consistency Has a Computable Blind Spot: A Commutation Theory of Label-Free Reliability for Vision-Language Figure Reading
- 基于对易性理论,识别输入扰动下不变性失效的隐藏错误
- 构建可计算的等变-一致性评分,无需训练即可检测模型缺陷
- 适用于评估多模态模型鲁棒性,尤其适合无标注数据场景
视觉语言模型的无标签可靠性依赖于输入扰动下的答案不变性。然而,存在一种系统性误读,其在扰动下仍保持不变,被误判为正确。本文证明该错误是可计算的:当扰动与错误可对易时,错误对模型不可见。随着扰动集增加,不可见错误集合(联合中心化子)逐渐缩小,可显式写出而非猜测。我们转而采用等变性思路:修改图像数据后,正确答案应以可计算方式改变。一对匹配的编辑可完全覆盖仿射读取错误;仅交换编辑无法覆盖标签置换错误,但循环重标记可弥补大部分缺口。据此构建了无需训练的等变-一致性评分,并发布 REND-EQUIV,对相同数据配对使用不变性和等变性编辑集。预测排序在三种模型和一个不受选择环路影响的手标数据集上均成立;另一不变性方法验证盲点属关系本质,非实现问题;循环重标记在真实样本上实现预期增益。该理论同样解释了分类器元测试文献中报告的排序反转现象:可检测性是关系与故障类别的共同属性,而非关系本身决定。
原文摘要 · Abstract (English)
Label-free reliability for vision-language models rests on invariance: perturb the input and a faithful reader's answer should not change. This has a known blind spot, a systematic misreading survives the perturbation and gets certified wrong, which we show is computable, not just real: an error is invisible to an edit exactly when the two commute, so the errors a suite cannot reach form its joint centralizer, a set that shrinks as edits are added and can be written down rather than guessed at. We act on the complementary relation, equivariance: edit a figure's data and the correct answer must change by a computable amount. Two matched edits are provably complete for affine reading errors; no suite of swap edits is complete for label permutations, and cyclic relabeling closes most of that gap. We instantiate the theory as the Equivariance-Consistency Score, a label-free, training-free detector, and release REND-EQUIV, pairing matched invariance and equivariance sets over identical data. The predicted ordering holds across three models and a hand-labeled population immune to the one circularity in how it is selected; a second invariance-family method confirms the blind spot belongs to the relation, not to any implementation; and cyclic relabeling delivers its predicted gain on a matched real sample. The same characterization explains a reported inversion of this ordering in the classifier metamorphic-testing literature: detectability is a joint property of the relation and the fault class, never of the relation alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。