arXiv:2505.13061cs.CV2025-05NeurIPS被引 3

研究发现机器也会被3D视觉错觉骗,提出新方法提升深度估计准确率。

3D Visual Illusion Depth Estimation

  • 利用视觉语言模型常识,自适应融合双目与单目深度信息。
  • 在近3000场景、20万图像数据上测试,现有方法均被错觉误导。
  • 新框架在多种错觉场景下表现最优,适合高精度深度感知任务。

3D视觉错觉是一种感知现象,通过操纵二维平面来模拟三维空间关系,使平面作品或物体在人眼视觉系统中呈现立体感。本文揭示,机器视觉系统同样会被3D视觉错觉严重误导,包括单目和双目深度估计。为探究和分析3D视觉错觉对深度估计的影响,我们构建了一个包含近3000个场景和20万张图像的大规模数据集,用于训练和评估当前最先进的单目与双目深度估计方法。同时,我们提出一种3D视觉错觉深度估计框架,利用视觉语言模型的常识知识,自适应融合双目视差与单目深度信息。实验表明,现有的最先进单目、双目及多视角深度估计方法均被各类3D视觉错觉所误导,而我们的方法在多个基准上达到了最先进性能。

原文摘要 · Abstract (English)

3D visual illusion is a perceptual phenomenon where a two-dimensional plane is manipulated to simulate three-dimensional spatial relationships, making a flat artwork or object look three-dimensional in the human visual system. In this paper, we reveal that the machine visual system is also seriously fooled by 3D visual illusions, including monocular and binocular depth estimation. In order to explore and analyze the impact of 3D visual illusion on depth estimation, we collect a large dataset containing almost 3k scenes and 200k images to train and evaluate SOTA monocular and binocular depth estimation methods. We also propose a 3D visual illusion depth estimation framework that uses common sense from the vision language model to adaptively fuse depth from binocular disparity and monocular depth. Experiments show that SOTA monocular, binocular, and multi-view depth estimation approaches are all fooled by various 3D visual illusions, while our method achieves SOTA performance.

深度估计视觉错觉多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。