用视频直接生成事件体素,150倍降存储,训练数据量大增。
V2V: Scaling Event-Based Vision through Efficient Video-to-Voxel Simulation
- 将视频帧直接转为体素网格,跳过高耗存储的事件流生成
- 存储减少150倍,用52小时1万段视频训练模型性能显著提升
- 适合想用海量数据训练事件视觉模型的研究者
事件相机具有高时间分辨率、高动态范围和低功耗等优势。然而,现有合成数据生成管道存储需求大、输入/输出负担重,真实数据稀缺,导致事件视觉训练数据难以规模化,限制了模型发展与泛化能力。为此,我们提出视频到体素(V2V)方法,可直接将常规视频帧转换为事件体素网格表示,完全绕过耗存储的事件流生成过程。V2V实现存储需求降低150倍,同时支持运行时参数随机化,增强模型鲁棒性。利用该效率,我们在总计52小时、涵盖10,000段多样视频上训练多个视频重建与光流估计模型架构,数据规模较现有事件数据集扩大一个数量级,取得显著性能提升。
原文摘要 · Abstract (English)
Event-based cameras offer unique advantages such as high temporal resolution, high dynamic range, and low power consumption. However, the massive storage requirements and I/O burdens of existing synthetic data generation pipelines and the scarcity of real data prevent event-based training datasets from scaling up, limiting the development and generalization capabilities of event vision models. To address this challenge, we introduce Video-to-Voxel (V2V), an approach that directly converts conventional video frames into event-based voxel grid representations, bypassing the storage-intensive event stream generation entirely. V2V enables a 150 times reduction in storage requirements while supporting on-the-fly parameter randomization for enhanced model robustness. Leveraging this efficiency, we train several video reconstruction and optical flow estimation model architectures on 10,000 diverse videos totaling 52 hours--an order of magnitude larger than existing event datasets, yielding substantial improvements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。