arXiv:2412.19067cs.CVcs.LG2024-12被引 4

利用运动补偿提升事件相机单目深度估计精度

Learning Monocular Depth from Events via Egomotion Compensation

  • 通过物理运动原理建模,使深度推测更可解释
  • 在真实与合成数据集上误差降低最高达10%
  • 适合高动态、低光照等极端场景下的深度感知研究

事件相机是受类脑启发的传感器,稀疏且异步地报告亮度变化。其高时间分辨率、高动态范围和低功耗特性使其适用于解决单目深度估计中的挑战(如高速或低光照条件)。然而,现有方法通常将事件流视为黑箱学习系统,未融合先验物理规律,导致模型过度参数化,未能充分利用事件数据中的丰富时序信息。为此,我们引入物理运动原理,提出一种可解释的单目深度估计框架,通过运动补偿效应显式评估不同深度假设的可能性。为此,我们设计了焦点代价判别(FCD)模块,以边缘清晰度为焦点水平的关键指标,并结合空间上下文辅助代价估计。此外,我们分析了框架内的噪声模式,提出新的跨假设代价聚合(IHCA)模块,通过代价趋势预测和多尺度一致性约束优化代价体积。在真实世界和合成数据集上的大量实验表明,所提框架在绝对相对误差指标上优于前沿方法最高达10%,展现出更优的预测精度。

原文摘要 · Abstract (English)

Event cameras are neuromorphically inspired sensors that sparsely and asynchronously report brightness changes. Their unique characteristics of high temporal resolution, high dynamic range, and low power consumption make them well-suited for addressing challenges in monocular depth estimation (e.g., high-speed or low-lighting conditions). However, current existing methods primarily treat event streams as black-box learning systems without incorporating prior physical principles, thus becoming over-parameterized and failing to fully exploit the rich temporal information inherent in event camera data. To address this limitation, we incorporate physical motion principles to propose an interpretable monocular depth estimation framework, where the likelihood of various depth hypotheses is explicitly determined by the effect of motion compensation. To achieve this, we propose a Focus Cost Discrimination (FCD) module that measures the clarity of edges as an essential indicator of focus level and integrates spatial surroundings to facilitate cost estimation. Furthermore, we analyze the noise patterns within our framework and improve it with the newly introduced Inter-Hypotheses Cost Aggregation (IHCA) module, where the cost volume is refined through cost trend prediction and multi-scale cost consistency constraints. Extensive experiments on real-world and synthetic datasets demonstrate that our proposed framework outperforms cutting-edge methods by up to 10\% in terms of the absolute relative error metric, revealing superior performance in predicting accuracy.

事件相机深度估计运动补偿可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。