arXiv:2606.23078eess.IV2026-06综述被引 1

系统梳理事件相机表征学习方法,揭示其如何平衡精度与效率。

A Systematic Survey on Event Camera Representation Learning

论文配图:A Systematic Survey on Event Camera Representation Learning
图 1 · 摘自论文原文
  • 按单/多表征分类,区分密集化与稀疏化处理思路。
  • 总结主流基准与任务评估设置,覆盖高低层视觉任务。
  • 指出自适应表征优化等未来方向,适合感知系统研究者。

事件相机具有微秒级延迟和高动态范围的优势,适用于挑战性感知任务。受生物视觉启发,它们输出异步稀疏事件流而非密集图像帧,与主流神经网络存在根本不匹配。本综述从将原始事件流转化为可学习表征的角度,系统回顾了近期进展。我们首先根据是否依赖单一核心表征或联合利用多种互补表征进行分类。单表征方法分为基于密集化的结构化网格表示(保留空间时间规律性)和基于稀疏化的原生离散时空结构表示(保持事件稀疏特性)。多表征方法则分为密集-密集与密集-稀疏混合形式,以融合不同表征优势。该表征中心的分类体系阐明了各类范式在结构规则性、时间保真度、稀疏性保持与架构兼容性之间的权衡。针对每种范式,我们分析其设计选择、建模原理及任务层面影响。此外,我们总结了代表性高层感知与底层视觉任务的标准基准与评估设置。最后,讨论开放问题,展望未来方向:从固定表征设计转向自适应表征优化,提升保真度与效率的平衡,构建更可扩展的事件感知系统。

原文摘要 · Abstract (English)

Event cameras offer distinctive advantages, including microsecond-level latency and high dynamic range, rendering them promising for challenging perception tasks. Inspired by biological vision, they output asynchronous and sparse event streams rather than dense image frames, creating a fundamental mismatch with mainstream neural networks. This survey reviews recent advances in event camera representation learning from the perspective of converting raw event streams into learnable representations. We first organize existing methods according to whether they rely on a single principal event representation or jointly exploit multiple complementary representations. Single-representation methods are further categorized into dense-based representations, which regularize events into structured grid-like forms, and sparse-based representations, which preserve event-native discrete spatio-temporal structures. Multi-representation methods are organized into dense-dense and dense-sparse hybrid formulations that exploit complementary representation properties. This representation-centric taxonomy clarifies how different paradigms balance structural regularity, temporal fidelity, sparsity preservation, and architectural compatibility. For each paradigm, we examine the underlying design choices, modeling principles, and task-level implications. We further summarize standard benchmarks and evaluation settings across representative high-level perception and low-level vision tasks. Finally, we discuss open problems and outline future directions from fixed representation design toward adaptive representation optimization, improved fidelity-efficiency trade-offs, and more scalable event-based perception systems.

事件相机表征学习视觉感知神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。