发现视觉模型无法识别错误边缘,导致深度预测严重崩溃。
Geometric Collapse: When Vision Models Fail to Verify Physical Causality

- 设计'打乱边缘'实验,制造违反物理规律的假边缘
- 错误边缘使预测偏差最大达干净图像的3.2倍
- 即使知道错误位置,修复效果也仅47%,适合研究模型可靠性
大规模自监督学习虽提升了密集几何预测能力,但其在推理时是否具备物理合理性判断仍不明确。本文提出'打乱边缘'(Scrambled Edges)这一受控反事实方法,在保持高频能量与结构匹配的前提下,注入违反表面连续性、光照一致性及遮挡顺序的显著边缘线索。在NYU Depth v2和KITTI数据集上,对各类CNN/ViT/SSL深度预测器测试发现,该方法引起的预测偏差比能量匹配噪声大最多3.2倍;扩散模型与流匹配模型虽响应减弱,但仍存在显著崩溃。这种几何崩溃具有全局传播性:即便已知污染区域,输出层修复仅能恢复47%精度,且误差广泛扩散至掩码外。结果表明当前密集预测器缺乏有效机制隔离物理不成立的边缘线索,亟需引入显式合理性评分与选择性特征融合。
原文摘要 · Abstract (English)
Recent progress in large-scale self-supervised learning has improved dense geometric prediction, but it remains unclear whether such scaling yields inference-time physical plausibility checks. We propose Scrambled Edges, a controlled counterfactual that injects salient edge-like cues while violating surface continuity, illumination coherence, and occlusion ordering. With energy-matched and structure-matched controls, we isolate the effect of unsupported edge evidence from high-frequency energy and edge sparsity. Across CNN/ViT/SSL depth predictors on NYU Depth v2 and KITTI, Scrambled Edges induce up to 3.2x larger deviation from clean predictions than energy-matched noise; additional diffusion and flow-matching depth estimators show attenuated but still significant collapse. The resulting Geometric Collapse propagates globally: even with oracle knowledge of the corrupted region, output-level repair recovers only 47%, with substantial error outside the mask. These findings provide controlled behavioral evidence that current dense predictors lack reliable mechanisms to quarantine physically unsupported edge cues, motivating explicit plausibility scoring and selective cue integration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。