arXiv:2509.15224cs.CV2025-09ICCV被引 10

用视觉大模型生成事件相机深度伪标签,无需昂贵标注即可实现高精度单目深度估计。

Depth AnyEvent: A Cross-Modal Distillation Paradigm for Event-Based Monocular Depth Estimation

  • 通过跨模态蒸馏,利用对齐的事件与RGB数据生成密集深度伪标签。
  • 在真实和合成数据上表现媲美全监督方法,且无需昂贵深度标注。
  • 基于深度任意模型(DAv2)改进的递归结构,达到当前最优性能。

事件相机以稀疏、高时间分辨率的方式捕捉视觉信息,特别适合高速运动和光照剧烈变化的环境。然而,缺乏大规模带有密集真值深度标注的数据集,制约了基于学习的单目深度估计在事件数据上的发展。为此,我们提出一种跨模态蒸馏范式,利用视觉基础模型(VFM)生成密集的代理标签。该策略只需空间对齐的事件流与RGB帧,设置简单,甚至可直接使用现成设备,并充分利用大规模VFM的鲁棒性。此外,我们提出适配VFM的方法,包括直接使用如Depth Anything v2(DAv2)等通用模型,或从中衍生出一种新型循环架构,用于从单目事件相机推断深度。我们在合成与真实世界数据集上评估该方法,结果表明:(i) 本跨模态范式在无需昂贵深度标注的情况下,性能媲美全监督方法;(ii) 基于VFM的模型达到当前最佳水平。

原文摘要 · Abstract (English)

Event cameras capture sparse, high-temporal-resolution visual information, making them particularly suitable for challenging environments with high-speed motion and strongly varying lighting conditions. However, the lack of large datasets with dense ground-truth depth annotations hinders learning-based monocular depth estimation from event data. To address this limitation, we propose a cross-modal distillation paradigm to generate dense proxy labels leveraging a Vision Foundation Model (VFM). Our strategy requires an event stream spatially aligned with RGB frames, a simple setup even available off-the-shelf, and exploits the robustness of large-scale VFMs. Additionally, we propose to adapt VFMs, either a vanilla one like Depth Anything v2 (DAv2), or deriving from it a novel recurrent architecture to infer depth from monocular event cameras. We evaluate our approach with synthetic and real-world datasets, demonstrating that i) our cross-modal paradigm achieves competitive performance compared to fully supervised methods without requiring expensive depth annotations, and ii) our VFM-based models achieve state-of-the-art performance.

事件相机深度估计跨模态蒸馏视觉大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。