视觉证据价值随观察状态变化,动态重排更利于跨场景决策。
Conditional Visual Evidence Utility: State-Dependent Rank Reversals in Frozen Vision-Language Encoders

- 在控制条件下逐项暴露颜色、形状、纹理证据,测量其条件边际效用。
- 800场景验证中,冻结的OpenCLIP与SigLIP出现显著状态依赖的排序反转。
- 动态重排可提升后续决策性能,尤其在查询与评估模式不一致时。
静态重要性评分将视觉证据压缩为单一排序,但观察一个线索后,剩余证据的价值可能改变。我们在受控的组合式视觉搜索任务中研究这一现象,可独立暴露颜色、形状和纹理证据,并在不同获取状态下测量其条件边际效用。在800个场景的保留验证中,冻结的OpenCLIP和SigLIP均表现出稳健的状态依赖排序反转,且集中出现在设计用于引发排序变化的候选重叠区域。该结构在两种证据累积方式和十种等价查询表述下持续存在,但在查询-场景错位时消失。我们还考察这些反转是否影响决策:在事后探索性匹配首行动分析中,仅在首次获取后重新排序,当决策在一种证据模式、表述或主干模型下选择,而在另一种下评估时,仍能获得正向第二步效用。结果表明,在此受控设置中,证据重要性具有状态依赖性,且更新证据排序可在评估者变化时保持决策相关价值。这推动了基于条件而非单一静态排名来评估视觉语言证据使用,也为未来自适应证据选择方法提供了可度量目标。
原文摘要 · Abstract (English)
Static importance scores compress visual evidence into a single ranking, but the value of remaining evidence can change after one cue has been observed. We study this possibility in controlled compositional visual search, where color, shape, and texture evidence can be independently exposed and their conditional marginal utility measured across acquisition states. In a held-out confirmation on 800 scenes, frozen OpenCLIP and SigLIP exhibit robust state-dependent rank reversals that concentrate in candidate-overlap regimes designed to induce ordering changes. The structure persists across two evidence-accumulation constructions and ten equivalent query wordings, but disappears under query-scene derangement. We also ask whether these reversals matter for decisions. In a post-confirmation exploratory matched-first-action analysis, reranking only after the first acquisition yields positive step-2 utility when decisions are selected under one evidence mode, wording, or backbone and evaluated under another. Together, these results show that evidence importance is state-dependent in this controlled setup and that updating an evidence ordering can retain decision-relevant value across evaluator changes. They motivate evaluating vision-language evidence use conditionally rather than through a single static ranking, while providing a measurable target for future adaptive evidence-selection methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。