arXiv:2511.02953cs.CVcs.AI2025-11被引 1

构建130亿事件的大规模事件相机数据集,提升深度估计泛化能力

EvtSlowTV -- A Large and Diverse Dataset for Event-Based Depth Estimation

  • 从YouTube视频构建超大规模事件数据集,覆盖多场景运动
  • 130亿事件数据支持自监督学习,提升复杂场景泛化性能
  • 无需帧标注,保留事件数据异步特性,适合真实环境应用

事件相机具有高动态范围(HDR)和低延迟特性,为复杂环境下的鲁棒深度估计提供了新可能。然而,现有事件基深度估计方法受限于小规模标注数据集,难以推广至真实场景。为此,我们提出EvtSlowTV,一个从公开YouTube视频中构建的大规模事件相机数据集,涵盖超过130亿事件,覆盖季节性徒步、飞行、风景驾驶及水下探索等多种环境与运动模式。该数据集规模较现有同类数据集大一个数量级,提供无约束的自然场景用于事件基深度学习。实验证明,该数据集适用于自监督学习框架,可充分挖掘原始事件流的HDR潜力;使用其训练的模型显著提升了在复杂场景与运动中的泛化能力。该方法无需帧级标注,保持了事件数据的异步特性。

原文摘要 · Abstract (English)

Event cameras, with their high dynamic range (HDR) and low latency, offer a promising alternative for robust depth estimation in challenging environments. However, many event-based depth estimation approaches are constrained by small-scale annotated datasets, limiting their generalizability to real-world scenarios. To bridge this gap, we introduce EvtSlowTV, a large-scale event camera dataset curated from publicly available YouTube footage, which contains more than 13B events across various environmental conditions and motions, including seasonal hiking, flying, scenic driving, and underwater exploration. EvtSlowTV is an order of magnitude larger than existing event datasets, providing an unconstrained, naturalistic setting for event-based depth learning. This work shows the suitability of EvtSlowTV for a self-supervised learning framework to capitalise on the HDR potential of raw event streams. We further demonstrate that training with EvtSlowTV enhances the model's ability to generalise to complex scenes and motions. Our approach removes the need for frame-based annotations and preserves the asynchronous nature of event data.

事件相机深度估计自监督学习大数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。