用CARLA模拟器检测交通物体,发现仿真数据训练的模型在真实场景下表现差。
How Real is CARLAs Dynamic Vision Sensor? A Study on the Sim-to-Real Gap in Traffic Object Detection
- 用CARLA的动态视觉传感器生成仿真事件数据训练模型
- 纯仿真训练的模型在真实数据上性能下降超30%
- 强调需改进仿真精度和领域自适应技术
事件相机因其低延迟、高时间分辨率和能效优势,正被广泛应用于交通监控。然而,事件感知检测模型的开发受限于真实世界标注数据集的稀缺。为解决此问题,已有多个仿真工具用于生成合成事件数据,其中CARLA驾驶模拟器内置了动态视觉传感器(DVS)模块以模拟事件相机输出。尽管如此,事件感知检测中的仿真到现实差距仍缺乏系统研究。本文通过仅使用CARLA的DVS生成的合成数据训练循环视觉变换器模型,并在不同比例的合成与真实事件流上测试其表现。实验表明,仅在合成数据上训练的模型在合成数据占比高的测试集上表现良好,但随着真实数据比例上升,性能显著下降;而基于真实数据训练的模型则展现出更强的跨域泛化能力。本研究首次对基于CARLA DVS的事件感知检测中仿真到现实的差距进行了量化分析,揭示了当前DVS仿真保真度的局限性,强调了在神经形态视觉用于交通监控时,亟需提升领域自适应技术。
原文摘要 · Abstract (English)
Event cameras are gaining traction in traffic monitoring applications due to their low latency, high temporal resolution, and energy efficiency, which makes them well-suited for real-time object detection at traffic intersections. However, the development of robust event-based detection models is hindered by the limited availability of annotated real-world datasets. To address this, several simulation tools have been developed to generate synthetic event data. Among these, the CARLA driving simulator includes a built-in dynamic vision sensor (DVS) module that emulates event camera output. Despite its potential, the sim-to-real gap for event-based object detection remains insufficiently studied. In this work, we present a systematic evaluation of this gap by training a recurrent vision transformer model exclusively on synthetic data generated using CARLAs DVS and testing it on varying combinations of synthetic and real-world event streams. Our experiments show that models trained solely on synthetic data perform well on synthetic-heavy test sets but suffer significant performance degradation as the proportion of real-world data increases. In contrast, models trained on real-world data demonstrate stronger generalization across domains. This study offers the first quantifiable analysis of the sim-to-real gap in event-based object detection using CARLAs DVS. Our findings highlight limitations in current DVS simulation fidelity and underscore the need for improved domain adaptation techniques in neuromorphic vision for traffic monitoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。