用事件相机提升暗光视频语义分割,兼顾精度与效率
Event-guided Low-light Video Semantic Segmentation
- 融合事件相机的运动信息引导图像分割
- 在3个数据集上达到领先性能,参数量仅1/11
- 适合低光照场景下的实时视频分割应用
近期视频语义分割方法在光照充足环境下表现良好,但在暗光条件下因可视度低、上下文信息不足,性能显著下降。同时,暗光环境加剧了帧间时序不一致性,导致视频闪烁。相比传统摄像头,事件相机能捕捉运动动态、过滤时间冗余信息,且对光照变化鲁棒。为此,我们提出EVSNet,一种轻量级框架,利用事件模态引导统一的光照不变表示学习。具体地,设计运动提取模块,从事件模态中提取短时与长时运动信息;运动融合模块自适应整合图像特征与运动特征;时序解码器则利用视频上下文生成分割预测。EVSNet架构轻量却实现领先性能。在3个大规模数据集上的实验表明,其优于现有方法,参数效率最高提升11倍。
原文摘要 · Abstract (English)
Recent video semantic segmentation (VSS) methods have demonstrated promising results in well-lit environments. However, their performance significantly drops in low-light scenarios due to limited visibility and reduced contextual details. In addition, unfavorable low-light conditions make it harder to incorporate temporal consistency across video frames and thus, lead to video flickering effects. Compared with conventional cameras, event cameras can capture motion dynamics, filter out temporal-redundant information, and are robust to lighting conditions. To this end, we propose EVSNet, a lightweight framework that leverages event modality to guide the learning of a unified illumination-invariant representation. Specifically, we leverage a Motion Extraction Module to extract short-term and long-term temporal motions from event modality and a Motion Fusion Module to integrate image features and motion features adaptively. Furthermore, we use a Temporal Decoder to exploit video contexts and generate segmentation predictions. Such designs in EVSNet result in a lightweight architecture while achieving SOTA performance. Experimental results on 3 large-scale datasets demonstrate our proposed EVSNet outperforms SOTA methods with up to 11x higher parameter efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。