用未来事件预测构建高效事件相机表示,实时处理高动态场景。
Fast Feature Field ($\text{F}^3$): A Predictive Representation of Events
- 基于未来事件预测构建稀疏事件的多通道时空图像表示
- 在HD分辨率下达120Hz,VGA下440Hz,支持多平台多环境应用
- 适合需要低延迟视觉感知的机器人系统,如自动驾驶与无人机
本文提出一种名为快速特征场(F³)的事件相机数据表示方法,通过从历史事件预测未来事件来学习表征,有效保留了场景结构与运动信息。F³利用事件数据的稀疏性,对噪声和事件率变化具有鲁棒性,结合多分辨率哈希编码与深度集合思想,实现高效计算:在高清(HD)分辨率下达到120 Hz,VGA分辨率下达440 Hz。F³将连续时空体积内的事件编码为多通道图像,支持多种下游任务。在三个机器人平台(汽车、四足机器人、飞行器)上,涵盖昼夜、室内外、城市及非铺装路面等多种场景与动态视觉传感器(不同分辨率与事件率),均取得光学流估计、语义分割与单目度量深度估计的领先性能。相关实现可在高清分辨率下以25–75 Hz完成任务预测。
原文摘要 · Abstract (English)
This paper develops a mathematical argument and algorithms for building representations of data from event-based cameras, that we call Fast Feature Field ($\text{F}^3$). We learn this representation by predicting future events from past events and show that it preserves scene structure and motion information. $\text{F}^3$ exploits the sparsity of event data and is robust to noise and variations in event rates. It can be computed efficiently using ideas from multi-resolution hash encoding and deep sets - achieving 120 Hz at HD and 440 Hz at VGA resolutions. $\text{F}^3$ represents events within a contiguous spatiotemporal volume as a multi-channel image, enabling a range of downstream tasks. We obtain state-of-the-art performance on optical flow estimation, semantic segmentation, and monocular metric depth estimation, on data from three robotic platforms (a car, a quadruped robot and a flying platform), across different lighting conditions (daytime, nighttime), environments (indoors, outdoors, urban, as well as off-road) and dynamic vision sensors (resolutions and event rates). Our implementations can predict these tasks at 25-75 Hz at HD resolution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。