arXiv:2412.06708cs.CVcs.RO2024-12NeurIPS被引 11

让事件相机在不同频率下都能精准检测物体,适应动态环境。

FlexEvent: Towards Flexible Event-Frame Object Detection at Varying Operational Frequencies

  • 用自适应融合模块整合高频事件与RGB语义信息
  • 支持20到180Hz的频率变化,最高达180Hz仍保持准确
  • 适合需要实时感知的机器人、自动驾驶等场景

事件相机凭借微秒级时间分辨率和异步工作模式,在动态环境中具有显著优势。然而,现有事件检测方法受限于固定频率范式,难以充分发挥事件数据的高时序分辨率与可适应性。为此,我们提出FlexEvent框架,实现多频率下的目标检测。该方法包含两个核心组件:FlexFuse——一种自适应事件帧融合模块,将高频事件数据与RGB帧中的丰富语义信息融合;FlexTune——一种频率自适应微调机制,生成适配不同频率的标签,提升模型在不同运行频率下的泛化能力。实验表明,该方法在大规模事件相机数据集上超越现有最先进方法,在标准与高频设置下均取得显著提升。尤其在20 Hz至90 Hz间动态切换时表现稳健,最高可在180 Hz下实现精确检测,验证了其在极端条件下的有效性。本框架为事件感知目标检测树立新基准,推动更灵活、实时的视觉系统发展。

原文摘要 · Abstract (English)

Event cameras offer unparalleled advantages for real-time perception in dynamic environments, thanks to the microsecond-level temporal resolution and asynchronous operation. Existing event detectors, however, are limited by fixed-frequency paradigms and fail to fully exploit the high-temporal resolution and adaptability of event data. To address these limitations, we propose FlexEvent, a novel framework that enables detection at varying frequencies. Our approach consists of two key components: FlexFuse, an adaptive event-frame fusion module that integrates high-frequency event data with rich semantic information from RGB frames, and FlexTune, a frequency-adaptive fine-tuning mechanism that generates frequency-adjusted labels to enhance model generalization across varying operational frequencies. This combination allows our method to detect objects with high accuracy in both fast-moving and static scenarios, while adapting to dynamic environments. Extensive experiments on large-scale event camera datasets demonstrate that our approach surpasses state-of-the-art methods, achieving significant improvements in both standard and high-frequency settings. Notably, our method maintains robust performance when scaling from 20 Hz to 90 Hz and delivers accurate detection up to 180 Hz, proving its effectiveness in extreme conditions. Our framework sets a new benchmark for event-based object detection and paves the way for more adaptable, real-time vision systems.

事件相机目标检测实时感知自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。