发现单目深度模型在模糊输入下会生成虚假3D结构,提出检测与修复方案。
The 3D Mirage: Probing and Taming 3D Hallucinations
- 构建首个结合上下文变化的真实世界幻觉测试集3D-Mirage。
- 设计双指标评估体系,量化3D结构偏差与上下文不稳定性。
- 提出轻量级自蒸馏方法,在不遗忘背景知识前提下消除幻觉。
单目深度基础模型通过学习大规模语义先验实现优异泛化能力,但由此引发关键缺陷:在平面或低曲率但感知模糊的输入上会生成虚假3D结构,我们称之为3D幻觉。本文提出端到端框架,用于探测、评估并缓解这一未被充分重视的安全风险。为探测,提出3D-Mirage,首个结合上下文变化与精确标注的真实世界幻觉基准,支持真实物体排除、多表面建模,专为压力测试单目深度在真实幻觉下的表现而设计。为评估,提出基于二阶幅度的评价方法,包含偏差综合得分(DCS)衡量高阶3D结构偏差,以及混淆综合得分(CCS)衡量上下文不稳定性。为缓解,引入接地自蒸馏策略,在Depth-Anything-V2基线基础上,精准定位并修复幻觉区域,同时保留背景知识,避免灾难性遗忘。本工作建立创新诊断与修复流程,呼吁将单目深度估计的评估从像素级精度转向结构与上下文鲁棒性。
原文摘要 · Abstract (English)
Monocular depth foundation models achieve remarkable generalization by learning large-scale semantic priors, but this creates a critical vulnerability: they hallucinate illusory 3D structures from planar/low-curvature but perceptually ambiguous inputs. We term this failure the 3D Mirage. This paper introduces a novel end-to-end framework to probe, score, and tame this under-quantified safety risk in monocular depth under context variation. To probe, we present 3D-Mirage, the first benchmark to combine context variation and precise annotation for real-world illusions with real object exclusions, multi-surface support; purpose-built to stress-test monocular depth on real-world illusions. To score, we propose a second-order magnitude-based evaluation with two metrics: the Deviation Composite Score (DCS) for high second-order 3D structure and the Confusion Composite Score (CCS) for contextual instability. To tame this failure, we introduce Grounded Self-Distillation, a parameter-efficient strategy on Depth-Anything-V2 baseline that surgically targets and resolves hallucination on illusion ROIs while preserving background knowledge, avoiding catastrophic forgetting. Our work provides an innovative pipeline for diagnosing and addressing this phenomenon, urging a necessary shift in the evaluation of MDE from pixel-wise accuracy to structural and contextual robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。