arXiv:2412.03259cs.CV2024-12被引 1

构建可控几何变换的事件数据生成器,助力事件视觉模型研究

GERD: Geometric event response data generation

  • 设计GERD模拟器,精确控制仿射、伽利略和时间缩放变换
  • 支持三类噪声模型与亚像素运动,生成带真实变换标签的数据
  • 适合研究事件视觉中几何不变性,推动模型可解释性

事件视觉传感器具备高时序分辨率、高动态范围和低功耗优势,但事件视觉模型性能仍落后于传统帧基视觉方法。我们认为这一差距部分源于对支配事件流的变换群缺乏系统研究。受几何与群论在计算机视觉中推动进展的启发,本文提出GERD:一个用于生成物体在精确控制的仿射、伽利略和时间缩放变换下的事件记录的模拟器。通过在每个时间步提供真实变换标签,GERD使几何特性在真实数据集或现有模拟器中难以分离的情况下,仍可进行假设驱动的受控研究。GERD支持三种噪声模型及亚像素运动,作为真实传感器数据集的补充。我们通过在文献模型上进行带有几何保证的训练评估展示了其应用价值,并将GERD开源发布。

原文摘要 · Abstract (English)

Event-based vision sensors offer high temporal resolution, high dynamic range, and low power consumption, yet event-based vision models lag behind conventional frame-based vision methods. We argue that this gap is partly due to the lack of principled study of the transformation groups that govern event-based visual streams. Motivated by the role that geometric and group-theoretic methods have played in advancing computer vision, we present GERD: a simulator for generating event-based recordings of objects under precisely controlled affine, Galilean, and temporal scaling transformations. By providing ground-truth transformations at each timestep, GERD enables hypothesis-driven and controlled studies of geometric properties that are otherwise hard to isolate in real-world datasets or with current event simulators. GERD supports three noise models and sub-pixel motion as a complement to real sensor datasets. We demonstrate its use in training by evaluating models from the literature with geometric guarantees and release GERD as an open tool available at

事件视觉数据生成几何变换仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。