让深度估计像人眼一样主动聚焦,看清透光场景的多层结构。
DepthFocus: Controllable Depth Estimation for See-Through Scenes
- 通过可调节的视觉变压器动态控制深度感知
- 在复杂透光场景中实现领先精度,显著优于现有方法
- 适合需要精准三维感知的自动驾驶与机器人领域
现实世界的深度往往不是单一的。透明材质会产生多层模糊,困扰传统感知系统。现有模型多为被动式,通常仅估计最近表面的静态深度图,即使近期的多头扩展也受限于固定特征表示。这与人类视觉主动调焦以感知目标深度的能力形成对比。本文提出深度聚焦(DepthFocus),一种可调控的视觉变压器,将立体深度估计重构为条件感知控制。模型不提取固定特征,而是根据物理参考深度动态调节计算,引入双重条件机制,选择性地感知与目标焦点对齐的几何结构。基于新构建的大规模合成数据集,DepthFocus 在所有评估基准上均达到最先进水平,涵盖单层与复杂多层场景。在保持不透明区域高精度的同时,能有效解决透明和反射场景中的深度歧义,有选择地重建目标距离的几何结构。该能力使感知更具意图驱动性,显著优于现有多层方法,标志着向主动三维感知迈出重要一步。
原文摘要 · Abstract (English)
Depth in the real world is rarely singular. Transmissive materials create layered ambiguities that confound conventional perception systems. Existing models remain passive; conventional approaches typically estimate static depth maps anchored to the nearest surface, and even recent multi-head extensions suffer from a representational bottleneck due to fixed feature representations. This stands in contrast to human vision, which actively shifts focus to perceive a desired depth. We introduce \textbf{DepthFocus}, a steerable Vision Transformer that redefines stereo depth estimation as condition-aware control. Instead of extracting fixed features, our model dynamically modulates its computation based on a physical reference depth, integrating dual conditional mechanisms to selectively perceive geometry aligned with the desired focus. Leveraging a newly curated large-scale synthetic dataset, \textbf{DepthFocus} achieves state-of-the-art results across all evaluated benchmarks, including both standard single-layer and complex multi-layered scenarios. While maintaining high precision in opaque regions, our approach effectively resolves depth ambiguities in transparent and reflective scenes by selectively reconstructing geometry at a target distance. This capability enables robust, intent-driven perception that significantly outperforms existing multi-layer methods, marking a substantial step toward active 3D perception. \noindent \textbf{Project page}: \href{https://junhong-3dv.github.io/depthfocus-project/}{\textbf{this https URL}}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。