提出新方法解析冻结视觉编码器的组合性,避免结果混淆。
Support Operation Factorization: Compositional Readout of Frozen Vision Encoders under Controlled Interventions

- 设计注入式评估协议,分离支持与操作因子
- 在Shapes3D和COCO上达到0.874和0.799的精确度
- 揭示不同模型在标签学习下的性能边界
对冻结视觉编码器进行组合分析,需同时判断什么变化以及何处变化。标准因子探测器分别评分,但会奖励重复使用相同预测槽的操作,称为‘操作洗白’。本文引入支持-操作网格上的注入对齐留一单元评估协议,及SO-OPF读出机制,将单元能量分解为支持显著性与竞争性操作后验。该方法区分两个被聚合评分混淆的问题:已知网格时载体是否组合未包含的绑定;以及能否从扁平单元标签恢复该网格。使用冻结的DINOv3特征,在Shapes3D-Extended上实现0.874的注入精度,在全局图像不重叠的COCO上为0.799;从扁平标签学习分配时,分别为0.769和0.762。在Shapes3D上采用匹配轴感知监督,因子化载体使学习分配精度从0.653提升至0.841,消除其洗白差距。SigLIP2在COCO上复现了该分离效果。重建的MuJoCo基底暴露边界:使用DINOv3时学习分配精度为0.569,使用SigLIP2时为0.484,出现显著槽坍缩。因此,因子化读出与注入评估能恢复两个基底中的未包含绑定,同时揭示而非隐藏渲染器特异性失败边界;它们并未建立从扁平标签的通用恢复能力。
原文摘要 · Abstract (English)
Compositional analysis of frozen vision encoders should determine both what changed and where it changed. Standard factor probes score these axes separately, however, and can reward multiple operations that reuse the same predicted slot. We call this failure operation laundering. We introduce an injectively aligned leave-one-cell-out protocol over support x operation grids and SO-OPF, a readout that factors cell energy into support salience and a competitive operation posterior. This formulation separates two questions that aggregate scores conflate: whether the carrier composes held-out bindings when the grid is known, and whether that grid can be recovered from flat cell labels. With frozen DINOv3 features, known factorial assignment reaches 0.874 injective accuracy on Shapes3D-Extended and 0.799 on globally image-disjoint COCO; learning the assignment from flat labels reaches 0.769 and 0.762, respectively. Under matched-axis-aware supervision on Shapes3D, the factored carrier improves learned-assignment accuracy from 0.653 to 0.841 over a dense carrier and eliminates its laundering gap. SigLIP2 replicates the COCO separation. A rebuilt MuJoCo substrate exposes a boundary: learned-assignment accuracy is 0.569 with DINOv3 and 0.484 with SigLIP2, with substantial slot collapse. Thus factored readout and injective evaluation recover held-out bindings on two substrates while exposing, rather than hiding, a renderer-specific failure boundary; they do not establish universal recovery from flat labels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。